Paper deep dive
Understanding Strategic Platform Entry and Seller Exploration: A Stackelberg Model
Garrett Seo, Xintong Wang, David C. Parkes
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/22/2026, 5:01:46 AM
Summary
The paper presents a Stackelberg game-theoretic model to analyze the strategic interaction between digital platforms and third-party sellers. It examines how platform entry policies—specifically the timing of entry into product markets—influence seller innovation and exploration behavior. Using a Gittins-index policy for single-seller scenarios and deep reinforcement learning for multi-seller environments, the authors demonstrate that optimal platform entry timing can balance revenue maximization with the preservation of market diversity and consumer welfare.
Entities (5)
Relation Signals (3)
Platform → implements → Entry Policy
confidence 95% · the platform acts as the leader by committing to an entry policy
Seller → uses → Gittins-index policy
confidence 95% · We characterize the seller's optimal explore-exploit strategy via a Gittins-index policy
Platform Entry → impacts → Seller Innovation
confidence 90% · Our findings highlight the incentives that drive platform entry and seller innovation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Online market platforms play an increasingly powerful role in the economy. An empirical phenomenon is that platforms, such as Amazon, Apple, and DoorDash, also enter their own marketplaces, imitating successful products developed by third-party sellers. We formulate a Stackelberg model, where the platform acts as the leader by committing to an entry policy: when will it enter and compete on a product? We study this model through a theoretical and computational framework. We begin with a single seller, and consider different kinds of policies for entry. We characterize the seller's optimal explore-exploit strategy via a Gittins-index policy, and give an algorithm to compute the platform's optimal entry policy. We then consider multiple sellers, to account for competition and information spillover. Here, the Gittins-index characterization fails, and we employ deep reinforcement learning to examine seller equilibrium behavior. Our findings highlight the incentives that drive platform entry and seller innovation, consistent with empirical evidence from markets such as Amazon and Google Play, with implications for regulatory efforts to preserve innovation and market diversity.
Tags
Links
- Source: https://arxiv.org/abs/2603.14206v1
- Canonical: https://arxiv.org/abs/2603.14206v1
Trouble viewing inline? Open PDF directly →
Full Text
73,305 characters extracted from source content.
Expand or collapse full text
by Understanding Strategic Platform Entry and Seller Exploration: A Stackelberg Model Garrett Seo Rutgers UniversityNew Brunswick, NJUSA garrett.seo@rutgers.edu , Xintong Wang Rutgers UniversityNew Brunswick, NJUSA xintong.wang@rutgers.edu and David C. Parkes Harvard UniversityCambridge, MAUSA parkes@eecs.harvard.edu (2026) Abstract. Online market platforms play an increasingly powerful role in the economy. An empirical phenomenon is that platforms, such as Amazon, Apple, and DoorDash, also enter their own marketplaces, imitating successful products developed by third-party sellers. We formulate a Stackelberg model, where the platform acts as the leader by committing to an entry policy: when will it enter and compete on a product? We study this model through a theoretical and computational framework. We begin with a single seller, and consider different kinds of policies for entry. We characterize the seller’s optimal explore-exploit strategy via a Gittins-index policy, and give an algorithm to compute the platform’s optimal entry policy. We then consider multiple sellers, to account for competition and information spillover. Here, the Gittins-index characterization fails, and we employ deep reinforcement learning to examine seller equilibrium behavior. Our findings highlight the incentives that drive platform entry and seller innovation, consistent with empirical evidence from markets such as Amazon and Google Play, with implications for regulatory efforts to preserve innovation and market diversity. Platform economy, Gittins index, reinforcement learning, agent-based modeling, multi-agent simulation This research is funded in part by Rutgers SAS Research Grant in Academic Themes. Our code is available at https://github.com/chailab-rutgers/platform-entry. †journalyear: 2026†copyright: c†conference: Proceedings of the ACM Web Conference 2026; April 13–17, 2026; Dubai, United Arab Emirates†booktitle: Proceedings of the ACM Web Conference 2026 (W ’26), April 13–17, 2026, Dubai, United Arab Emirates†doi: 10.1145/3774904.3792678†isbn: 979-8-4007-2307-0/2026/04†ccs: Computing methodologies Multi-agent systems 1. Introduction Modern digital platforms like Amazon, Apple’s App Store, and DoorDash have become dominant intermediaries for economic transactions, with their success built on third-party vendors and developers (“sellers”) who bring diverse services to the market. These platforms also increasingly operate as direct competitors to sellers, by launching their own products. Examples include AmazonBasics, Apple’s own apps in its App store, and the “DoorDash Kitchens” which emerged after the pandemic-driven food delivery boom. The digital era has intensified this competition, bringing new advantages to platforms. With unparalleled access to data, platforms can identify and imitate popular products at lower risk and cost than sellers (Zhu and Liu, 2018). Furthermore, platforms have the power to promote their own items in search and recommendations, and operate at a scale unmatched by typically resource-constrained sellers. The impact of platform entry on products is highly contested. It may benefit consumers (buyers) with lower prices and better quality, but it may also discourage third-party innovation if sellers expect the platform to imitate their products and take their profits. This contentious dynamic has also become a focus of global antitrust enforcement. In a landmark 2023 lawsuit, the U.S. Federal Trade Commission, together with 17 states, sued Amazon, alleging that the company uses non-public seller data to illegally maintain its monopoly power by targeting lucrative markets for its own private-label products (Federal Trade Commission, 2023). This was followed in March 2024 by a U.S. Department of Justice lawsuit against Apple for allegedly suppressing innovative apps and technologies that could weaken the iPhone’s dominance (U.S. Department of Justice, 2024). The European Union’s Digital Markets Act (DMA), which became fully applicable in 2024, restricts platforms from using non-public business data to gain a competitive advantage (European Parliament, Council of the European Union, 2022). This landscape raises critical questions for both platform operators and regulators: How should a platform decide whether and when to enter a product market, in balancing its own commercial interests against the health of its seller ecosystem? Do a platform’s private objectives align with social welfare, or is regulatory intervention necessary? To answer these questions, we develop a game-theoretic model of the strategic interaction between a platform and its sellers. We model the platform-seller relationship as a Stackelberg game, where the platform, acting as the leader, first decides its policy as to when to enter and compete on each of one or more products. This positioning as the leader in the game reflects a platform’s commitment power (e.g., through published fees or policies and exclusivity clauses). The sellers, after observing the platform’s policy, then decide on their product innovation and sales strategy. Our Contributions. We first study a single-seller model, where the seller’s decision-making is modeled as a costly search problem. By successfully adapting the Gittins index to this setting, we derive a closed-form expression for the seller’s value of exploring an untested product (“exploration value”) which incorporates the platform’s entry policy, allowing for a precise characterization of the seller’s optimal explore/exploit policy. We show that the platform’s revenue is piecewise monotonic with respect to the entry time or fee that parametrizes its policy, and develop an algorithm to optimize the platform’s policy by identifying a finite set of Pareto optimal ones. We demonstrate this algorithm on three kinds of platform policies: a global entry policy that applies a universal entry time across products, a global entry combined with transaction fees policy, and a heterogeneous entry policy for product-specific entry times. We then generalize our model to a multi-seller environment to reflect the complex interactions on real-world platforms. As direct analytical solutions become intractable, we develop a multi-agent simulator and use deep reinforcement learning to find approximate Nash equilibria in the sellers’ game. This allows us to solve for the platform’s optimal entry policy under a range of distinct, strategically motivated environments (“clustered” and “diverse”), analyzing how platform entry affects different market scenarios. The single and multi-seller simulations reveal several insights on the platform’s strategic role. We find that a rational platform relies on seller exploration to discover high-demand products, which often increases overall consumer welfare (or social welfare) and encourages product exploration when viable alternative products exist, relative to no platform entry. The platform’s optimal entry is market-dependent: in “clustered markets” with a dominant product, a rational platform encourages seller diversification. The optimal timing, however, can be sensitive to the product’s risk profile, e.g., high-stakes products require a longer protection period to incentivize seller innovation. In “diverse markets” with specialized sellers, an aggressive (early) entry policy can be destructive, forcing sellers to abandon their niches and cluster around selling safer products, reducing overall market diversity and welfare. This dynamic compels a rational platform to commit to delay its entry, demonstrating that while its private incentives are not perfectly aligned with social welfare, they also prevent very bad outcomes. These findings suggest that, for a sufficiently forward-looking platform, the objective of revenue maximization is largely aligned with broader measures of market ecosystem health, such as product diversity and consumer welfare. The challenge, though, is to find the market-specific, non-myopic policies that a platform can use to achieve this alignment. This work provides a framework for understanding this strategic trade-off, offering a methodology for platform designers to optimize for long-term growth, and a lens for regulators to assess the market impact of platform competition. 1.1. Related Work Empirical Studies of Platform Entry. There are several empirical work documenting platform entry and competition with their third-party sellers. Amazon, rather than avoiding competition, tends to enter product spaces of high demand that are offered by many vendors (Zhu and Liu, 2018). Such entry decision increases the likelihood of third-party sellers exiting the platform. In the mobile app market, the threat of platform entry can cause developers to divert innovation efforts towards niche product categories to establish early competitive positions (Wen and Zhu, 2019). Our work gives a game-theoretic model to provide counterfactual analysis of the strategic incentives behind these empirical observations. Economic Models of Hybrid Platforms. Prior literature has developed economic models of hybrid platforms to analyze the trade-offs of their dual role (Anderson and Bedre Defolie, 2024), the impact of data regulations (Madsen and Vellodi, 2025), and specific strategies like self-preferencing in search rankings (Hagiu et al., 2022) and the use of vertical control mechanisms (Kang and Muir, 2022). We contribute to this line of work by offering a computational model focused specifically on how a platform’s entry policy affects seller innovation, product diversity, and market efficiency. Contract Design for Exploration. The seller’s problem in our model can be viewed as a variant of costly sequential search, rooted in the classic “Pandora’s Box” problem (Weitzman, 1979) and multi-armed bandit frameworks (Gittins and Jones, 1979). Related work studies how a principal can guide an agent’s search using explicit contracts, such as linear payment schemes (Dütting et al., 2023) or formal delegation mechanisms (Bechtel et al., 2022; Hoefer et al., 2025; Ivanov et al., 2024). In our model, the platform’s entry policy provides an implicit contract, shaping the exploration behavior of strategic sellers. RL for Economic Platform Design. Our approach aligns with recent work using RL and simulation to design and understand complex economic systems, including dynamically setting reserve prices in auctions (Shen et al., 2020), selling user impressions to advertisers (Tang, 2017), designing tax policies (Zheng et al., 2022), optimizing user satisfaction for recommender systems (Chen et al., 2019; Zhan et al., 2021), and designing sequential price mechanisms (Brero et al., 2021). Consistent with some of this literature, we use a Stackelberg game to model the platform-user interaction, a framework previously applied to analyze strategic problems like collusion mitigation (Brero et al., 2022) and platform fee-setting under market shocks (Wang et al., 2023). Gittins Index. The multi-armed bandit (MAB) problem models a sequential decision problem where an agent balances exploring less well-understood options with exploiting better known ones to maximize cumulative reward. For a specific class of MAB problems, the Gittins index provides a provably optimal solution (Gittins, 1974). This powerful result allows a complex, multi-dimensional optimization problem to be decomposed into a series of independent calculations, one for each arm. At each decision point, the optimal policy is to simply choose the arm with the highest current Gittins index. We demonstrate the applicability of the Gittins index to the seller problem in the single-seller setting. We defer details on the definition and assumptions to Appendix A. 2. A Single-Seller Model with Platform Entry We first consider the interaction between a platform and a single third-party seller, and analyze how the platform’s policy influences the seller’s exploratory behavior. 2.1. A Stackelberg Model of Platform Entry 2.1.1. Shared Product Space. We consider a set of M potential products, each initially in an undeveloped (U) state. The demand or reward for each product is unknown on the platform before being sold. We assume the platform only enters products after a seller has explored a product and revealed its demand (e.g., from public sales ranking of products). 2.1.2. The Seller’s Problem. To explore an undeveloped product j, the seller incurs a one-time innovation cost cjc_j. Upon exploration, the product’s demand is revealed: it transitions to a “good” state (G) with reward rjgr^g_j and probability pjp_j or a “bad” state (B) with reward rjbr^b_j and probability 1−pj1-p_j. We assume rjg>rjb≥0r^g_j>r^b_j≥ 0 and that the seller has accurate priors of pjp_j, rjgr^g_j, and rjbr^b_j, informed by sources such as market research, historical data, or similar products. These parameters capture essential differences in product types, reflecting variations in development costs and market positioning (e.g., luxury versus practical, niche versus mass market). The innovation cost can also reflect a seller’s expertise within a product domain. The seller has a capacity-constraint of offering one product at a time, and seeks to maximize their total expected discounted reward, given the platform’s policy. 2.1.3. The Platform’s Policy. The platform commits to a policy _p that includes an entry-time parameter TpjT_p_j as part of its strategic design. This parameter specifies the number of timesteps the platform waits before entering a developed product j that has been revealed to be in its “good” state by the seller, transitioning product j to an entered (E) state. This reflects data patents, or empirical observations that platforms such as Amazon typically enter only after products demonstrate strong sales and positive reviews (Zhu and Liu, 2018). For simplicity, we assume that once the platform enters, it captures the entire reward stream rjgr^g_j. The platform can further introduce a fraction parameter α, representing the reward split between the platform and the seller. The parameter α can be interpreted as a transaction fee that applies regardless of the product’s realized state. Given these components, in Section 2.3, we analyze three platform policy settings: a global entry TpT_p for all products, a global entry TpT_p with transaction fee α, and a heterogeneous entry =(Tp1,Tp2,…,TpM)T_p=(T_p_1,T_p_2,…,T_p_M) allowing distinct entry times for each product. 2.2. Seller’s Optimal Policy Under _p Without platform entry, the seller faces a multi-arm bandit problem where the product opportunities are independent stochastic processes. In this setting, the Gittins-index policy provides a provably optimal strategy (Gittins and Jones, 1979). We now introduce platform policy and denote the resulting problem by ℳM. The platform policy modifies the state-transition dynamics, so we model each product as a Markov chain with state space Sj=U,G,B,ES_j=\U,G,B,E\, representing the undeveloped, good, bad, and entered by platform states, respectively. We derive a closed-form, piece-wise expression for the Gittins index that incorporates the seller’s optimal stopping decision in response to a platform entry TpjT_p_j. For each product j, we calculate the Gittins index by identifying the optimal stopping time. We consider the optimal stopping rule by partitioning the state space into a stopping set S and a continuation set =Sj∖C=S_j . If the product is in a state s∈s , the seller continues to sell j; if s∈s , the seller stops selling j In our platform entry model, it suffices to consider the following three different stopping rules, with their corresponding indices. (1) Stopping rule τ1 _1: =US=\U\. The seller does not explore. (2) Stopping rule τ2 _2: =ES=\E\. The seller should continue to gain rewards from a bad state, and stop at platform entry. (3) Stopping rule τ3 _3: =B,ES=\B,E\. The seller stops if the product is revealed to be bad or if the platform enters. We do not consider state Sj=GS_j=G to be in the stopping set as rjgr_j^g is the highest reward achievable in any given state. We defer the detailed closed-form derivation of these stopping-time indices under platform policy, _p to Appendix B.1. Let Gj(k)(U;)G_j^(k)(U; _p) denote the Gittins index associated with a specific stopping rule τk _k for the undeveloped product j. The Gittins index for product j is the maximum value across all stopping rules: Gj(U;)=maxGj(1)(U;),Gj(2)(U;),Gj(3)(U;)G_j(U; _p)= \G_j^(1)(U; _p),G_j^(2)(U; _p),G_j^(3)(U; _p) \ However, the optimality of the Gittins-index policy relies on independence across products, so we first check whether the independence among products is preserved, and under what conditions. Proposition 2.1. A platform’s policy _p preserves the products as independent stochastic processes, only if the seller acts optimally in response. The seller’s optimal policy is the Gittins-index policy. Proof. It suffices to prove the claim for heterogeneous entry times pT_p. A global policy TpT_p corresponds to the special case Tpj=TpT_p_j=T_p for all products j, while a transaction fee α uniformly rescales rewards by (1−α)(1-α). Hence the results for TpT_p and (Tp,α)(T_p,α) follow directly. The proof proceeds in four steps: (1) Construct a rested bandit ℳ′M from the current version ℳM. (2) Show Vπℳ′≥VπℳV_π^M ≥ V_π^M for any seller policy π. (3) Characterize a set of seller policies π′π where Vπℳ′=VπℳV_π^M =V_π^M. (4) Show the Gittins-index policy belongs to this set. Step 1 (Rested Conversion ℳ′M ). The seller’s problem ℳM is described by a market state (t)=(,)x(t)=(S,C). S denotes the product state vector, where each SjS_j takes values in the state space defined previously. C denotes the product counter vector where each Cj∈∅,Tpj,Tpj−1,…,2,1,0C_j∈\ ,T_p_j,T_p_j-1,…,2,1,0\ tracks the time left before the platform enters product j. The counter is inactive (Cj=∅C_j= ) when product j is in state U or B, and equals 0 when the product has entered state E. ℳM is not made up of independent stochastic products. In particular, suppose product j is in state xj(t)=(G,Cj),x_j(t)=(G,C_j), meaning it is in its good state and the platform will enter after CjC_j timesteps. If the seller instead sells a different product j′≠j ≠ j at time t, then the counter for product j still decreases, so its state updates to xj(t+1)=(G,Cj−1),if Cj>1,(E,0),if Cj=1.x_j(t+1)= cases(G,C_j-1),&if C_j>1,\\ (E,0),&if C_j=1. cases Thus, the state of product j changes despite selling another product j′j . To recover independence, we define a rested version of the problem, denoted ℳ′M . ℳ′M matches ℳM in all state transitions, except for state G. Specifically, in ℳ′M , if product j is in state xj(t)=(G,Cj)x_j(t)=(G,C_j) and the seller chooses some other product j′≠j ≠ j, then product j remains unchanged: xj(t+1)=(G,Cj).x_j(t+1)=(G,C_j). However, if the seller chooses product j at time t, then we keep the original transition: xj(t+1)=(G,Cj−1),if Cj>1,(E,0),if Cj=1.x_j(t+1)= cases(G,C_j-1),&if C_j>1,\\ (E,0),&if C_j=1. cases Step 2 (Dominance of ℳ′M ). We now show that the seller always makes more utility in ℳ′M . We define a trajectory τ=((0),a(0),(1),a(1),…)τ=(x(0),a(0),x(1),a(1),…) which is the sequence of market states and actions generated in ℳ′M or ℳM. For a seller policy π, we define the expected discounted utility from initial state (0)x(0) as Vπ((0))=τ∼pπ(⋅∣(0))[∑t=0∞γtR((t),π((t))].V_π (x(0) )=E_τ p_π (· (0) ) [ _t=0^∞γ^t\,R (x(t),π(x(t) ) ]. Here, γ∈(0,1)γ∈(0,1) is the discount factor and R((t),a(t))R (x(t),a(t) ) is the reward at time t. Further, the trajectory distribution pπp_π is defined as: pπ(τ∣(0))=∏t=0∞π(a(t)∣(t))((t+1)∣(t),a(t)),p_π(τ (0))= _t=0^∞π (a(t) (t) )\,T (x(t+1) (t),a(t) ), where T denotes the transition probability function, and pπp_π describes the probability of the seller playing some trajectory τ under policy π. To show that the expected reward of the seller in ℳ′M is greater than or equal to in ℳM, or that Vπℳ′((0))≥Vπℳ((0))V_π^M (x(0) )≥ V_π^M (x(0) ), we show that the reward from any given trajectory in ℳ′M is greater than that in ℳM. Fix an arbitrary trajectory τ, and consider any time t along this trajectory with state-action pair (ℳ′(t),a(t)) (x^M (t),a(t) ) and (ℳ(t),a(t)) (x^M(t),a(t) ) for ℳ′M and ℳM respectively. We claim that, for every action a(t)a(t) in the sequence, R(ℳ′(t),a(t))≥R(ℳ(t),a(t)).R (x^M (t),a(t) )\;≥\;R (x^M(t),a(t) ). We compare rewards across actions. Action a(t)=0a(t)=0. No product is sold, and the reward is 0 in both problems: R(ℳ′(t),a(t))=R(ℳ(t),a(t))=0.R (x^M (t),a(t) )=R (x^M(t),a(t) )=0. Action a(t)=ja(t)=j. The seller sells product j at time t. We distinguish three product states. Undeveloped. If xjℳ′(t)=(U,∅)x^M _j(t)=(U, ), then necessarily xjℳ(t)=(U,∅)x^M_j(t)=(U, ) as well. Therefore, R(ℳ′(t),a(t))=R(ℳ(t),a(t))=−cj+rjg,w.p. pj,rjb,w.p. 1−pj.R (x^M (t),a(t) )=R (x^M(t),a(t) )=-c_j+ casesr^g_j,&w.p. p_j,\\ r^b_j,&w.p. 1-p_j. cases Bad. If xjℳ′(t)=(B,∅)x^M _j(t)=(B, ), then xjℳ(t)=(B,∅)x^M_j(t)=(B, ) as well, and hence R(ℳ′(t),a(t))=R(ℳ(t),a(t))=rjb.R (x^M (t),a(t) )=R (x^M(t),a(t) )=r^b_j. Good. Suppose xjℳ′(t)=(G,Cjℳ′)x^M _j(t)=(G,C^M _j). Counters evolve differently across the two problems: in ℳ′M the counter decreases only when product j is selected, whereas in ℳM it decreases whenever the product is in the good state. Hence Cjℳ′≥CjℳC^M _j≥ C^M_j. If Cjℳ=0C^M_j=0, then xjℳ(t)=(E,0)x^M_j(t)=(E,0) while xjℳ′(t)=(G,Cjℳ′)x^M _j(t)=(G,C^M _j), so R(ℳ′(t),a(t))=rjg>R(ℳ(t),a(t))=0.R(x^M (t),a(t))=r^g_j>R(x^M(t),a(t))=0. Otherwise the product remains in state G in both ℳ′M and ℳM, and the rewards coincide, yielding Rℳ′=Rℳ=rjgR^M =R^M=r^g_j. Across all actions, the reward gained by the seller in ℳ′M is greater than or equal to the reward in ℳM at every time step along a trajectory τ. Because the two problems induce the same trajectory distribution under any fixed seller policy π (i.e., pπℳ′(τ∣0)=pπℳ(τ∣0)p^M _π(τ _0)=p^M_π(τ _0)), it follows that the expected discounted utility in ℳ′M is weakly higher, and Vπℳ′((0))≥Vπℳ((0))V_π^M (x(0) )≥ V_π^M (x(0) ). Step 3 (Policies Achieving Vπℳ′=VπℳV_π^M =V_π^M). We now describe the set of policies π′\π \ for which Vπℳ′((0))=Vπℳ((0))V_π^M (x(0) )=V_π^M (x(0) ). Previously, we showed that ℳ′M can yield a strictly greater reward than ℳM only when some product j is in the good state xjℳ′(t)=(G,Cjℳ′)x^M _j(t)=(G,C_j^M ), since the counter satisfies Cjℳ′≥CjℳC_j^M ≥ C_j^M. In order for both counters to be equivalent, we must have π′π satisfy that if xj(t)=(G,Cj)x_j(t)=(G,C_j), then π′((t))=jπ (x(t) )=j. In this way, the counter decreases the same way in both problems, and we can interpret π′π as a seller policy that always continues to sell a product in its G state until entry. Since ℳ′M is a MAB problem with independent stochastic arms, the Gittins-index policy π∗π^* is optimal (Weber, 1992). Step 4 (Gittins Optimality). To complete the proof, we must show that the Gittins-index policy π∗π^* belongs to the set of policies π′\π \. In order for π∗∈π′π^*∈\π \, π∗π^* satisfies that if xj(t)=(G,Cj)x_j(t)=(G,C_j), then π∗((t))=jπ^* (x(t) )=j. Suppose the seller follows π∗π^* and at time t product j is in the good state, xj(t)=(G,Cj)x_j(t)=(G,C_j). Since all products are initially undeveloped in (0)x(0), there exists some t′<t <t at which product j was first selected, with xj(t′)=(U,∅)x_j(t )=(U, ). By the Gittins-index policy, j maximized the index at time t′t , i.e., Gj(U;Tpj)≥Gj′∀j′≠j.G_j(U;T_p_j)≥ G_j ∀\,j ≠ j. Moreover, since Gj(G;Tpj)>Gj(U;Tpj),G_j(G;T_p_j)>G_j(U;T_p_j), product j continues to maximize the index at time t, implying π∗((t))=j.π^*(x(t))=j. ∎ The proof above highlights how platform policies shape optimal exploration incentives. In some market environments, however, exploration is effectively predetermined. We define a dominating product jdomj^dom as a product whose undeveloped Gittins index exceeds that of any other product under any platform policy: Gjdom(U;p)>Gj(U;p)∀j≠jdom,∀p.G_j^dom(U; π_p)>G_j(U; π_p) ∀ j≠ j^dom,\;∀ π_p. In experiments, we focus on market environments without dominating products, since sellers would always explore jdomj^dom first, limiting the influence of platform policies on exploration. 2.3. Finding the Optimal Platform Policy The platform’s objective is to choose the optimal ∗ π^*_p, that maximizes its own expected discounted utility, anticipating the seller’s optimal response. Instead of exhaustively searching over an infinite set of platform policies _p, we show that the policy space can be partitioned into regions, each with a distinct optimal seller strategy. We further find that the platform’s optimal policy lies among a finite set of candidate points across these regions. We first show that the platform’s utility, upu_p, is a piecewise function of its policy _p, with branches defined by the optimal seller strategy πs∗π^*_s that follows from _p. For this section, we simplify the market state as =(,t)x=(S,t), where S is the product state vector with each Sj∈U,B,G,ES_j∈\U,B,G,E\ indicating the state of product j, and t denotes the timestep. The utility for the platform can be described recursively: Up(,t;P)=0,j∗=∅,Fbad-cont(rj∗b,t,P),Sj∗=B,pj∗b(Fbad(rj∗b,t,P)+Up((Sj∗←B),t+1;P))+pj∗g(Fgood(rj∗g,t,P)+Up((Sj∗←E),t+Tpj;P))Sj∗=U U_p(S,t; π_P)= where j∗=πs∗(,t,P)=argmaxjGj(Sj,t;P)j^*=π^*_s(S,t, π_P)= _jG_j(S_j,t; π_P), or equivalently, the best product to sell at state (,t)(S,t) by the Gittins index. Fbad-contF_bad-cont, FbadF_bad, and FgoodF_good are closed-form expressions of the platform’s reward that depend on the policy setting. Fbad-contF_bad-cont represents the ongoing reward that the platform obtains when the seller continues to sell a product that is realized in its bad state for the remainder of the game. FbadF_bad and FgoodF_good correspond to the reward the platform receives when the seller’s product transitions to its bad or good state, respectively. We identify where the seller’s optimal policy may change with _p by defining a boundary set ℬB, made up of three types: (1) Zero boundary: bj0=Gj(U;)=0b_j^0=G_j(U; _p)=0. The seller is indifferent between exploring and not exploring product j. (2) Bad-indifference boundary: bj,j′B=Gj(B;)−Gj′(U;)=0b_j,j ^B=G_j(B; _p)-G_j (U; _p)=0. The seller is indifferent between selling product j, which has been realized in its bad state, and exploring an undeveloped product j′j . (3) Undeveloped-indifference boundary: bj,j′U=Gj(U;)−Gj′(U;)=0b_j,j ^U=G_j(U; _p)-G_j (U; _p)=0. The seller is indifferent between exploring product j and j′j . Each boundary b∈ℬb partitions the policy space _p into the sets :b()>0\ _p:b( _p)>0\ and :b()<0\ _p:b( _p)<0\. We defined a region as the non-empty intersection of such sets across all boundaries b∈ℬb , representing a subset of the platform policy space where the seller’s optimal choice of j∗j^* remains identical for a given market state (,t)(S,t). Thus, each piece of up()u_p( _p) corresponds to a region with a fixed seller strategy. Let ℛ(ℬ)=R1,R2,…,RkR(B)=\R_1,R_2,...,R_k\ be the collection of all nonempty regions induced by ℬB, where the seller strategy remains fixed for all ∈Ri _p∈ R_i. We can now define upu_p over all branches Ri∈ℛ(ℬ)R_i (B) as: up()=up(|R1)=Up(U,0;R1),if ∈R1,⋮up(|Rk)=Up(U,0;Rk),if ∈Rk,Ri∈ℛ(ℬ).u_p( _p)= casesu_p( _p|R_1)=U_p(S^U,0;R_1),&if _p∈ R_1,\\[6.00006pt] \;\; &\;\; \\[6.00006pt] u_p( _p|R_k)=U_p(S^U,0;R_k),&if _p∈ R_k, cases R_i (B). To maximize upu_p, we maximize over all regions. Naively, this can be done by iterating through all possible policies within a region. Instead, we show that for every region Ri∈ℛ(ℬ)R_i (B), up(|Ri)u_p( _p|R_i) is monotone with respect to each component of _p. Thus, maximizing a region reduces to maximizing over a set of Pareto optimal platform strategies within the region for which no other strategy improves all monotone components simultaneously. Given a policy setting, every monotonically increasing component (e.g., transaction fee) is bounded above and every monotonically decreasing component (e.g., entry time) is bounded below within RiR_i, so the set of Pareto optimal strategies is finite. Below we analyze three platform policy settings. 2.3.1. Global Entry A global entry policy is defined with a single platform entry time =Tp∈ℕ _p=T_p , the same for all products. Denote γ as the platform discount factor. The closed-form functions F, representing the platform’s reward, are: Fbad-cont(rjb,t,P=Tp) F_bad-cont(r^b_j,t, π_P=T_p) =Fbad(rjb,t,P=Tp)=0, =F_bad(r^b_j,t, π_P=T_p)=0, Fgood(rjg,t,P=Tp) F_good(r^g_j,t, π_P=T_p) =rjgγt+Tp1−γ. =r_j^g γ^t+T_p1-γ. FgoodF_good increases as TpT_p decreases, Fgood(rjg,t,Tp+1)<Fgood(rjg,t,Tp),∀Tp∈ℕ.F_good(r^g_j,t,T_p+1)<F_good(r^g_j,t,T_p),\;∀\;T_p . Since up(Tp)u_p(T_p) is a weighted sum of FgoodF_good, up(Tp)u_p(T_p) decreases monotonically with TpT_p. Further, TpT_p is bounded below, so every region RiR_i can be maximized over a finite set of Pareto optimal points. Because each region RiR_i is one-dimensional in TpT_p, identifying this set reduces to selecting the smallest TpT_p in that region. 2.3.2. Global Entry and Transaction Fee We extend the platform’s policy =(Tp,α) _p=(T_p,α) to include a transaction fee, α∈[0,1]α∈[0,1]. In this case, the closed-form functions F are: Fbad-cont(rjb,t,P=(Tp,α)) F_bad-cont(r^b_j,t, π_P=(T_p,α)) =αrjbγt1−γ, =α r_j^b γ^t1-γ, Fbad(rjb,t,P=(Tp,α)) F_bad(r^b_j,t, π_P=(T_p,α)) =αrjbγt, =α r_j^bγ^t, Fgood(rjg,t,P=(Tp,α)) F_good(r^g_j,t, π_P=(T_p,α)) =αrjgγt(1−γTp)1−γ+rjgγt+Tp1−γ. =α r_j^g γ^t(1-γ^T_p)1-γ+r_j^g γ^t+T_p1-γ. All F are linear and nonnegative in α, and thus increase monotonically with α. Meanwhile, FgoodF_good remains decreasing in TpT_p: Fgood(Tp+1)−Fgood(Tp) F_good(T_p+1)-F_good(T_p) =−rjgγt+Tp(1−α)1−γ≤0. =-r_j^g γ^t+T_p(1-α)1-γ≤ 0. Since γ∈(0,1)γ∈(0,1), rjg≥0r_j^g≥ 0, and α∈[0,1]α∈[0,1], the difference is always nonpositive. Thus, Fgood(Tp+1)≤Fgood(Tp)F_good(T_p+1)≤ F_good(T_p) for all Tp∈ℕT_p , and FgoodF_good decreases monotonically with TpT_p. upu_p is a weighted sum of the F functions, and therefore, increases monotonically with α and decreases monotonically with TpT_p within each region RiR_i. As established, TpT_p is bounded below, and α is also bounded above by 11, so every region RiR_i can be maximized over a finite set of Pareto optimal points. Formally, a policy (Tp,α)(T_p,α) is Pareto optimal if there exists no other point (Tp′,α′)(T_p ,α ) such that Tp′≤TpT_p ≤ T_p and α′≥α ≥α with at least one inequality strict. 2.3.3. Heterogeneous Entry We consider a flexible entry policy ==(Tp1,Tp2,…,TpM)∈ℕM _p=T_p=(T_p_1,T_p_2,…,T_p_M) ^M, where the platform can set an independent entry policy TpjT_p_j for each product j. The closed-form functions F are: Fbad-cont(rjb,t,P=) F_bad-cont(r^b_j,t, π_P=T_p) =Fbad(rjb,t,P=)=0, =F_bad(r^b_j,t, π_P=T_p)=0, Fgood(rjg,t,P=) F_good(r^g_j,t, π_P=T_p) =rjgγt+Tpj1−γ. =r_j^g γ^t+T_p_j1-γ. Here, FgoodF_good monotonically decreases with each TpjT_p_j and is bounded below, as shown in the global entry case. Thus, every region RiR_i can be maximized over a finite set of Pareto optimal points. Formally, a policy =(Tp1,Tp2,…,TpM)T_p=(T_p_1,T_p_2,…,T_p_M) is Pareto optimal if there exists no other ′=(Tp1′,Tp2′,…,TpM′)T _p=(T _p_1,T _p_2,…,T _p_M) such that Tpj′≤TpjT_p_j ≤ T_p_j for all j with at least one inequality strict. 2.3.4. Computation and Complexity We summarize the algorithm for computing the optimal platform policy in Appendix B.2. We provide intuition for the global entry and heterogeneous entry policy settings with toy examples in Appendix B.3 and B.4. The number of regions in the policy space depends on the setting. For global entry and global entry with transaction fees, the number of regions grows polynomially in the number of products due to the seller indifference boundaries. However, heterogeneous entry creates an M-dimensional policy space, causing the number of regions to grow exponentially with the number of products. Monotonicity and boundedness allow us to focus only on Pareto-optimal strategies within each region, despite this exponential complexity. 2.4. An Illustration of Platform Policy Choice In this section, we construct simple, representative market environments to study how a strategic platform may tailor its entry and fee policies according to market structure. We focus on identifying conditions under which platform incentives align with or diverge from seller exploration, and explore how regulatory interventions can mitigate platform’s excessive profit extraction. Table 1. Type A represents a product with moderate, stable payoffs and a relatively low cost, whereas Type B a riskier product with the potential of higher reward but higher cost. Type A Type B Cost 50 120 Reward 100, 50 200, 0 Probability 0.5, 0.5 0.2, 0.8 We use two types of products in Table 1 to construct environments reflecting different distributions of product opportunities: • A market with three Type A and one Type B products (3A1B), • A market with one Type A and three Type B products (1A3B). Table 2 presents the optimal platform policies and the associated utility and exploration metrics in the two markets. We set the seller’s discount factor to γs=0.9 _s=0.9, and assume a more forward-looking platform with γp=0.95 _p=0.95. We highlight the following: (1) Rational platform entry encourages exploration. This is reflected in the consistent increase in the number of products explored relative to the no-entry case. A rational platform avoids premature entry, since its profits remain partially aligned with seller exploration. (2) Market composition shapes optimal platform behavior. In markets dominated by safe products with predictable demand and low innovation costs (e.g., 3A1B), the platform optimally sets higher fees (40%) to reliably capture a steady revenue stream. This scenario likely reflects many real-world markets, driven by stable demand and incremental innovation. By contrast, in markets where demand is less predictable and success depends on innovating riskier products (e.g., 1A3B), the platform benefits from setting lower transaction fees (8%), better aligning its incentives with seller exploration and enabling it to imitate and monetize more new products. (3) High transaction fees reduce seller utility and exploration. In markets like 3A1B, the platform prioritizes profit extraction via higher fees over earlier entry to give sellers time to recoup from costs and explore later. This suppresses early exploration and buyer utility. Imposing caps on transaction fees limits the platform from extracting excessive profits while encouraging earlier entry, boosting product exploration and buyer utility. (4) Heterogeneous entry improves flexibility and buyer utility, but may hurt sellers. Allowing product-specific entry times enables the platform to imitate low-cost products earlier and high-cost products later, at the expense of seller profits. By placing a minimum entry barrier, we limit the platform from imitating too early, maintaining higher buyer utility while improving seller profit. Table 2. Agent utility and seller exploration metrics on markets with different product compositions. (a) Environment 3A1B Policy Setting Platform Seller Buyer Prod. Explored No platform entry or fee 0 865 2174 2.38 Tp∗=3T_p^*=3 2332 514 3447 3 (Tp∗,α∗)=(4,0.4)(T_p^*,α^*)=(4,0.4) 2633 292 3285 3 ∗=(3,3,3,8)T_p^*=(3,3,3,8) 2724 525 3967 4 Fee cap α≤0.2α≤ 0.2 (Tp∗,α∗)=(3,0.2)(T_p^*,α^*)=(3,0.2) 2555 386 3447 3 (b) Environment 1A3B Policy Setting Platform Seller Buyer Prod. Explored No platform entry or fee 0 876 2510 2.9 Tp∗=8T_p^*=8 1814 551 3098 4 (Tp∗,α∗)=(8,0.08)(T_p^*,α^*)=(8,0.08) 1921 485 3098 4 ∗=(1,7,7,7)T_p^*=(1,7,7,7) 2205 354 3259 4 Entry cap Tpj≥5T_p_j≥ 5 ∗=(5,8,8,8)T_p^*=(5,8,8,8) 2004 492 3217 4 Seller exploration is evaluated by expected products explored. Total buyer utility is the sum of discounted rewards from offered products, reflecting realized demand. Shaded rows indicate settings with a cap on transaction fees (left) or entry barrier (right). 3. A Multi-Seller Model with Platform Entry We now consider a platform with multiple sellers, capturing richer strategic dynamics, such as information spillover and market congestion. Here, the assumptions required for the optimality of the Gittins index for a seller’s problem no longer hold. Exploration by a seller now generates a public signal about the associated product, affecting the beliefs and strategies of all others. At the same time, as multiple sellers crowd into the same product space, individual payoffs may decline. Thus, the problem is no longer a set of independent search problems but a complex, non-stationary game. To study this more realisitic scenario, we extend our model and use multi-agent reinforcement learning to identify and analyze the resulting equilibria. We focus on the global entry policy setting, as solving for the optimal policy becomes more complex in the multi-seller case when multiple policy dimensions are involved. Our primary interest is to understand how sellers interact with each other under strategic platform entry. 3.1. The Multi-Seller Markov Game We model the strategic interaction as a Stackelberg game, where the platform commits to a global entry policy TpT_p and the sellers play a finite-horizon Markov Game ℳTpM_T_p induced by TpT_p. 3.1.1. The Sellers’ Game. The game ℳTpM_T_p is defined by: (1) Agents: A set of N sellers, indexed by i∈1,2,…,Ni∈\1,2,...,N\, each with a discount factor of γi _i. (2) Products: A set of M products, indexed by j∈1,2,…,Mj∈\1,2,...,M\. Each product has a known prior probability pjp_j of being “good” (high reward rjgr^g_j) or “bad” (low reward rjbr^b_j). Each seller i has a one-time innovation cost of ci,jc_i,j to explore product j. (3) State (xt∈x_t ): The state at time t includes (a) The state of each product: undeveloped, good, bad, entered, denoted by U,G,B,E\U,G,B,E\. (b) A N×MN× M matrix indicating which sellers are currently offering which products. (c) A N×MN× M matrix tracking the time elapsed since each seller first offered each product. (4) Actions (at∈a_t ): The joint action at=(a1,t,…,aN,t)a_t=(a_1,t,…,a_N,t) is the combination of individual seller actions, where ai,t∈0,1,2,…,Ma_i,t∈\0,1,2,...,M\ represents seller i’s choice of product j or nothing (i.e., 0). (5) Transitions (xt+1∼(xt,at)x_t+1 (x_t,a_t)): The state transitions based on the current state and the joint action. Product states change upon seller exploration and platform entry TpT_p. (6) Rewards (ri,tr_i,t): If seller i chooses action ai,t=ja_i,t=j, their reward is: ri,t=fj(r^j,nj,t)−ℐi,j⋅ci,j,r_i,t=f_j( r_j,n_j,t)-I_i,j· c_i,j, where r^j r_j is the realized reward of product j, nj,tn_j,t is the number of sellers offering product j, fjf_j is some function modeling the reward under different extent of market congestion, and ℐi,jI_i,j is an indicator variable that applies the cost on the first time seller i offers product j. If ai,t=0a_i,t=0 or product j has been entered by the platform, then ri,t=0r_i,t=0. (7) Observations (oi,t∈Ωo_i,t∈ ): The observation for seller sis_i at time t, oi,t⊂xto_i,t⊂ x_t, contains every information from xtx_t except for the private innovation cost of every other seller. 3.1.2. Seller and Platform Objectives. Each seller i chooses a policy πi _i to maximize their own total expected discounted reward. As each seller’s reward depends on the actions of other sellers, the goal is to find a policy profile ∗(Tp)=(π1∗,…,πN∗) π^*(T_p)=(π^*_1,…,π^*_N) that is a Nash equilibrium : πi∗∈argmaxπiπi,π−i∗[∑t=0Tγitri,t],∀i∈1,…,N. _i^*∈ _ _iE_ _i, _-i^* [ _t=0^T _i^tr_i,t ], ∀ i∈\1,...,N\. The platform’s objective is to choose an entry time Tp∗T^*_p that maximizes its own utility, anticipating the sellers’ equilibrium response π∗(Tp)π^*(T_p). The platform’s utility is the sum of discounted rewards from all products it has entered: uP(Tp)=π∗(Tp)[∑t=0T∑j=1Mγpt⋅rjg⋅ℐxj,t=E]u_P(T_p)=E_π^*(T_p) [ _t=0^T _j=1^M _p^t· r^g_j·I\x_j,t=E\ ] The platform’s optimization problem is thus: Tp∗=argmaxTp≥1uP(Tp).T_p^*= _T_p≥ 1u_P(T_p). 3.2. Approximating a Seller Equilibrium The introduction of multiple sellers competing with each other creates a non-stationary environment, making analytical solutions (including state space and joint action) appear intractable. Therefore, we use deep reinforcement learning to model the sellers’ strategic behavior. To address the non-stationarity of evolving seller strategies, we use an iterative best-response procedure to identify an approximate, ϵε-Nash equilibrium: (1) Independent training: Sellers are trained in parallel for a set of K episodes to learn policy profile, π=(π1,…,πN)π=( _1,…, _N). (2) Iterative best-response: For each agent, we evaluate the regret (or unilateral devation gain) by freezing the other agents’ policies π−i _-i and train a best-response policy πibr _i^br against them, starting from initially trained πi _i: Regreti(π)=max(0,πibr,π−i[∑t=0Tγtri,t]−πi,π−i[∑t=0Tγtri,t]) _i(π)= (0, _ _i^br, _-i [ _t=0^Tγ^tr_i,t ]-E_ _i, _-i [ _t=0^Tγ^tr_i,t ] ) (3) Convergence check: If the maximum regret across all agents falls below a threshold ϵε, the policy profile is an approximate, ϵε-Nash equilibrium, and the training terminates. Otherwise, we resume parallel training. While theoretical guarantees for MARL in general-sum games remain an open question, the iterative procedure enables the identification of empirically stable joint policies, corresponding to an approximate, ϵε-Nash equilibrium. To account for the possibility of multiple equilibria, we additionally execute this process with multiple random seeds to obtain a range of potential equilibrium outcomes for each TpT_p. 4. A Simulation Study of Strategic Platform Entry in Multi-Seller Markets We develop a multi-agent Gym simulation environment based on the multi-seller model in Section 3 to examine how strategic platform entry shapes market outcomes under different seller competition structures. The simulation allows us to explore a range of market configurations, facilitating counterfactual analysis and the interpretability of the market outcomes. 4.1. Experiment Settings We construct simulation environments that capture key strategic forces that can influence platform entry and seller exploration, while avoiding excessive model complexity. To this end, we focus on two canonical categories of market structures—clustered and diverse—which span the range of competitive conditions observed in real-world platforms. The clustered setting represents markets with a highly profitable or popular product space (e.g., sellers crowding into categories like tech accessories), whereas the diverse setting captures markets characterized by differentiated niches (e.g., merchants specializing in crafts like handmade jewelry or custom art). We adopt two representative sellers, capturing the simplest setting in which information spillover and competition can arise. Each seller can represent a group of similar sellers, allowing interpretable analysis of exploration and equilibrium under varying entry strategies with manageable computation. 4.1.1. Market Environment Configurations. While we cannot use the Gittins index to derive a seller solution in settings with multiple sellers, we employ a Gittins-index-guided design to ensure that sellers’ incentives align with the intended clustered or diverse market structures. Specifically, for each sampled product reward profile, we adjust the corresponding innovation costs so that sellers’ induced policies generate the desired market scenario in the absence of platform entry (i.e., Tp=∞T_p=∞). For all scenarios, we introduce “control products” with Gittins indices below a fixed threshold, G¯ G, to serve as background alternatives, ensuring that observed exploration behaviors arise endogenously from incentive structures rather than from a lack of available options. Below, we describe these distinct market environments and how they are generated. We denote Gi,jG_i,j as the formulation of the Gittins index of product j for seller i. • Clustered Environments There is a single popular product space j∗j^* whose shared-reward Gittins index exceeds the index of any other product j: Gi,j∗(U;Tp=∞)>G¯>Gi,j(U;Tp=∞), ∀i and j≠j∗.G_i,j^*(U;T_p=∞)> G>G_i,j(U;T_p=∞), $∀$i and j≠ j^*. We consider two scenarios within this category, analyzing how cost structure and reward size shape seller incentives under clustered competition. – Scenario C1 (Standard): A high-demand product space. – Scenario C2 (High-stakes): A high-stakes product space with a large innovation cost but also a higher reward. • Diverse Environments Sellers have specialized incentives, captured by their one-time innovation costs ci,jc_i,j, with each seller having a preferred product. We consider two scenarios, analyzing how asymmetries in seller capabilities shape strategic responses and product diversity. – Scenario D1 (Specialists): Each seller has a unique preferred product jij_i, and faces prohibitively high costs to explore other’s niche: Gi,ji(U;Tp=∞)>G¯>Gi,ji′(U;Tp=∞),for i≠i′ and ji≠ji′G_i,j_i(U;T_p=∞)> G>G_i,j_i (U;T_p=∞), i≠ i and j_i≠ j_i – Scenario D2 (Specialist and Generalist): This models a market with a specialist seller i∗i^* and a generalist seller i. The generalist can enter the specialist’s niche but still prefers their own product: Gi∗,ji∗(U;Tp=∞)>G¯>Gi∗,ji(U;Tp=∞),for ji∗≠jiG_i^*,j_i^*(U;T_p=∞)> G>G_i^*,j_i(U;T_p=∞), j_i^*≠ j_i Gi,ji(U;Tp=∞)>Gi,ji∗(U;Tp=∞)>G¯,for ji∗≠jiG_i,j_i(U;T_p=∞)>G_i,j_i^*(U;T_p=∞)> G, j_i^*≠ j_i To cover a variety of risk-reward profiles, we sample product parameters from discrete sets, specifically rjg∼75,100,200r^g_j \75,100,200\, rjb∼0,25,50r^b_j \0,25,50\, and pjg∼0.2,0.5,0.8p^g_j \0.2,0.5,0.8\.111In C2, we model high-stakes products by augmenting rjgr^g_j with an additional 500. 4.1.2. Agent Configuration. Given the large state space, we model each seller’s exploration using deep Q-learning (DQN), with best-response checks. Each seller trains its action–value function independently, and we iteratively compare policies to best responses against others’ fixed strategies to approximate an empirical ϵε-Nash equilibrium, using a regret threshold of ϵ=0.33ε=0.33 for convergence. To identify the optimal platform entry time, we search over integer values of TpT_p. For each value, we run the MARL training procedure under multiple random seeds to find seller equilibria and compute the platform’s expected revenue. We put all hyperparameter settings for training in Appendix C.1. 4.2. Effect of Platform Entry: Empirical Analysis Figure 1 presents our main findings on how platform entry affects different market structures, using several key metrics: • Agent utilities: total rewards for the platform, sellers, and consumers (i.e., rewards generated by the platform and sellers), • Products explored: the fraction of distinct products explored by sellers, measuring innovation, • Product variety: the fraction of distinct products offered per timestep, measuring product diversity, • Cluster rate: the frequency of sellers offering the same product. For a given TpT_p of an environment, we run a large number of simulations with different seeds on the resulting seller game under the trained DQN seller policy π∗π^* to capture these metrics. In case when there are multiple equilibria, we report a range of values for these equilibrium outcome metrics. (a) Standard cluster (C1) Utility Metrics (b) Standard cluster (C1) Product Metrics (c) High-stakes cluster (C2) Utility Metrics (d) High-stakes cluster (C2) Product Metrics (e) Diverse specialists (D1) Utility Metrics (f) Diverse specialists (D1) Product metrics (g) Diverse specialist and generalist (D2) Utility Metrics (h) Diverse specialist and generalist (D2) Product Metrics Figure 1. Expected utility metrics for the platform (blue), sellers (green), and buyers (yellow) as a function of TpT_p on left. Products explored (yellow), product variety (blue), and cluster rate (purple) as a function of TpT_p on right. These metrics are based on the results of 4,000 environment simulations. The range at each TpT_p denotes outcomes from multiple equilibria. In C1, optimal entry is early, Tp∗≈2T^*_p≈ 2, to capture value from the clustered product. In C2, optimal entry is delayed, Tp∗≈5T^*_p≈ 5, to cover high innovation costs. In D1, optimal entry, Tp∗≈2T^*_p≈ 2, has a muted effect on innovation. In D2, optimal entry is delayed, Tp∗≈11T^*_p≈ 11. 4.2.1. Clustered Environments. We see that a well-timed platform entry can promote seller diversification: the threat of entry on a high-demand product encourages sellers to explore alternatives. This echoes empirical evidence (Wen and Zhu, 2019), showing that app developers proactively shift innovation by improving their current apps or developing new apps before platform entry occurs (e.g., moving from general health apps to niche ski apps). In the high-stake cluster scenario C2 (Fig. 1(c)), the platform is incentivized to delay its entry until Tp∗=5T^*_p=5, capturing value from one seller exploring the high-demand product while prompting the other to innovate elsewhere. Here, the platform’s objective aligns more closely with seller exploration/diversification and social welfare, as sellers use the longer protection period to justify higher costs. Early entry can discourage the development of potentially lucrative products, limiting the platform’s value capture. In contrast, in the standard cluster scenario C1 (Fig. 1(a)), the optimal entry occurs earlier (Tp∗=2T^*_p=2) where the sellers still cluster despite early entry. This aggressive entry does not fully align with social welfare or market diversity goals, as the market favors a later entry (Tsw∗=8T^*_sw=8) to preserve seller incentives to explore riskier or lower-demand products, highlighting the potential need for regulatory intervention to discourage overly aggressive platform entry. 4.2.2. Diverse Environments. Overall, early platform entry either reduces market diversity (in D2) or has little impact relative to no-entry (in D1), whereas delayed entry increases market diversity. This effect is most pronounced in the specialist-generalist scenario D2 (Fig. 1(h)), which mirrors real-world online markets that feature a mix of niche innovators (specialists) and established sellers (generalists) that can pivot across product spaces. When facing early entry, the generalist seller may choose to sell clustered products preemptively, capturing short-term gains before imitation occurs, since their flexibility allows them to explore alternative products later (Fig. 1(h)). This mirrors empirical evidence (Wen and Zhu, 2019), which shows that app developers who have a portfolio of apps unaffected by entry can shift to existing apps more easily. This reduces overall product exploration, diversity, and buyer utility. Consequently, the platform’s optimal entry time is much later (Tp∗=11T^*_p=11), providing a longer protection period that encourages sellers to pursue niche space and maintain market diversity (Fig. 1(g)). In D1 (specialists) markets, by contrast, sellers occupy distinct niches. Platform entry can occur earlier without significant disruption, mainly to extract value rather than reshape seller behavior (Fig. 1(f)). 5. Discussion The optimal single-seller policy for a given platform policy _p, can be derived using a closed-form Gittins index. Given this, we solve for the optimal platform policy by optimizing over a finite set of Pareto-optimal points within regions corresponding to distinct seller strategies. In multi-seller settings, we use deep reinforcement learning to train seller policies and solve for approximately optimal platform entry under approximate seller equilibria. In the single-seller model, we explore different kinds of platform policies and consider different market compositions. Higher transaction fees are observed in markets with predictable demand and low innovation costs, while lower fees are observed in innovation-driven markets with uncertain demand. However, excessive fees reduce seller profit and slow innovation, making caps on fees effective in restoring seller profits while increasing buyer utility. Heterogeneous platform entry across products boosts buyer utility through flexible entry-timing but can hurt seller profits. Imposing limits on the earliest possible platform entry balances these effects. With multiple sellers, we explore seller-seller and platform-sellers interaction in simulation, across different, clustered and diverse, market environments. Platforms tend to enter aggressively when products are in high demand and offered by many sellers, but choose to commit to delayed entry for products with high costs or uncertain outcomes. This aligns with findings from Zhu and Liu (2018) that show Amazon enters categories like toys and games but avoids high-cost categories. More interestingly, sellers adapt even before entry occurs. Platform entry can disrupt clustered product markets, prompting some sellers to explore alternatives. Under aggressive platform entry, more versatile sellers may preemptively cluster into other sellers’ products to capture short-term gains before entry. Overall, our findings indicate that market structure plays a large role in determining whether platform entry aligns with seller exploration. In settings where alignment occurs, social welfare, innovation, and market diversity increases. When market structure favors aggressive entry or the platform enters early timing, seller exploration can be reduced compared to social-welfare optimal entry, and regulatory interventions may be needed to restore social welfare. References S. Anderson and Ö. Bedre Defolie (2024) Hybrid platform model: monopolistic competition and a dominant firm. The RAND Journal of Economics 55 (4), p. 684–718. Cited by: §1.1. C. Bechtel, S. Dughmi, and N. Patel (2022) Delegated pandora’s box. In Proceedings of the 23rd ACM Conference on Economics and Computation, p. 666–693. Cited by: §1.1. G. Brero, A. Eden, M. Gerstgrasser, D. C. Parkes, and D. Rheingans-Yoo (2021) Reinforcement learning of sequential price mechanisms. In Thirty-Fifth AAAI Conference on Artificial Intelligence, p. 5219–5227. Cited by: §1.1. G. Brero, E. Mibuari, N. Lepore, and D. C. Parkes (2022) Learning to mitigate ai collusion on economic platforms. In Proceedings of the 36th International Conference on Neural Information Processing Systems, Cited by: §1.1. M. Chen, A. Beutel, P. Covington, S. Jain, F. Belletti, and E. Chi (2019) Top-k off-policy correction for a reinforce recommender system. In Proceedings of the 12th ACM International Conference on Web Search and Data Mining., p. 456–464. Cited by: §1.1. P. Dütting, T. Ezra, M. Feldman, and T. Kesselheim (2023) Multi-agent contracts. In Proceedings of the 55th Annual ACM Symposium on Theory of Computing, STOC 2023, p. 1311–1324. Cited by: §1.1. European Parliament, Council of the European Union (2022) Regulation (eu) 2022/1925 of the european parliament and of the council of 14 september 2022 on contestable and fair markets in the digital sector and amending directives (eu) 2019/1937 and (eu) 2020/1828 (digital markets act). External Links: Link Cited by: §1. Federal Trade Commission (2023) Federal Trade Commission v. Amazon.com Inc. External Links: Link Cited by: §1. J. C. Gittins and D. M. Jones (1979) A dynamic allocation index for the discounted multiarmed bandit problem. Biometrika 66 (3), p. 561–565. Cited by: §1.1, §2.2. J. Gittins (1974) A dynamic allocation index for the sequential design of experiments. Progress in statistics, p. 241–266. Cited by: §1.1. A. Hagiu, T. Teh, and J. Wright (2022) Should platforms be allowed to sell on their own marketplaces?. The RAND Journal of Economics 53 (2), p. 297–327. Cited by: §1.1. M. Hoefer, C. Schecker, and K. Schewior (2025) Designing Exploration Contracts. In 42nd International Symposium on Theoretical Aspects of Computer Science (STACS 2025), Leibniz International Proceedings in Informatics (LIPIcs), Vol. 327, Dagstuhl, Germany, p. 50:1–50:19. External Links: ISBN 978-3-95977-365-2, ISSN 1868-8969, Link, Document Cited by: §1.1. D. Ivanov, P. Dütting, I. Talgam-Cohen, T. Wang, and D. C. Parkes (2024) Cited by: §1.1. Z. Y. Kang and E. Muir (2022) Contracting and vertical control by a dominant platform. In Proceedings of the 23rd ACM Conference on Economics and Computation, p. 694–695. Cited by: §1.1. E. Madsen and N. Vellodi (2025) Insider imitation. Journal of Political Economy 133 (2), p. 652–709. External Links: Document Cited by: §1.1. W. Shen, B. Peng, H. Liu, M. Zhang, R. Qian, Y. Hong, Z. Guo, Z. Ding, P. Lu, and P. Tang (2020) Reinforcement mechanism design: with applications to dynamic pricing in sponsored search auctions. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, p. 2236–2243. Cited by: §1.1. P. Tang (2017) Reinforcement mechanism design. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, p. 5146–5150. Cited by: §1.1. U.S. Department of Justice (2024) United states v. apple inc.. External Links: Link Cited by: §1. X. Wang, G. Q. Ma, A. Eden, C. Li, A. Trott, S. Zheng, and D. Parkes (2023) Platform behavior under market shocks: a simulation framework and reinforcement-learning based study. In Proceedings of the ACM Web Conference 2023, p. 3592–3602. Cited by: §1.1. R. Weber (1992) On the gittins index for multiarmed bandits. The Annals of Applied Probability 2 (4), p. 1024–1033. External Links: ISSN 10505164, Link Cited by: Appendix A, §2.2. M. L. Weitzman (1979) Optimal search for the best alternative. Econometrica 47 (3), p. 641–654. External Links: ISSN 00129682, 14680262 Cited by: §1.1. W. Wen and F. Zhu (2019) Threat of platform-owner entry and complementor responses: evidence from the mobile app market. Strategic Management Journal 40 (9), p. 1336–1367. Cited by: §1.1, §4.2.1, §4.2.2. R. Zhan, K. Christakopoulou, Y. Le, J. Ooi, M. Mladenov, A. Beutel, C. Boutilier, E. Chi, and M. Chen (2021) Towards content provider aware recommender systems: a simulation study on the interplay between user and provider utilities. In Proceedings of the Web Conference 2021, p. 3872–3883. Cited by: §1.1. S. Zheng, A. Trott, S. Srinivasa, D. C. Parkes, and R. Socher (2022) The AI economist: taxation policy design via two-level deep multiagent reinforcement learning. Science Advances 8 (18), p. eabk2607. Cited by: §1.1. F. Zhu and Q. Liu (2018) Competing with complementors: an empirical look at amazon.com. Strategic Management Journal 39 (10), p. 2618–2642. Cited by: §1.1, §1, §2.1.3, §5. Appendix A Gittins Index The optimality of the Gittins-index policy holds under the assumptions that each arm is an independent stochastic process with discounted rewards and is stationary, i.e., playing one arm does not influence the state or rewards of another and the state of an unplayed arm does not change over time (Weber, 1992). In our setting, each “arm” corresponds to a different product. The index, denoted Gj(xj)G_j(x_j), for an arm j in a given state xjx_j, is defined as the maximal expected reward rate, where the maximization is over all possible future stopping times τ≥1τ≥ 1: Gj(xj)=supτ≥1[∑t=0τ−1γtRj(xj(t))∣xj(0)=xj][∑t=0τ−1γt∣xj(0)=xj],G_j(x_j)= _τ≥ 1 E [ _t=0^τ-1γ^tR_j(x_j(t)) x_j(0)=x_j ]E [ _t=0^τ-1γ^t x_j(0)=x_j ], where γ is the discount factor, Rj(xj(t))R_j(x_j(t)) is the reward from arm j at time t, and the expectation [⋅]E[·] is taken over the stochastic evolution of the arm’s state. Appendix B Deferred Analysis for Section 2 B.1. Gittins Index Calculation We derive the Gittins index under each stopping rule. We generalize the derivation to include transaction fee α and for Tpj=TpT_p_j=T_p or Tpj∈T_p_j _p. In the global and heterogeneous entry case, set α=0α=0. We use V to denote the total expected discounted reward under a certain stopping rule, and D to denote the corresponding total expected discounted time horizon. Under τ1 _1, we have: Gj(1)(U;Tpj,α)=0G_j^(1)(U;T_p_j,α)=0 Under τ2 _2, we have: V(2) V^(2) =−cj+pj∑t=0Tpj−1γt(1−α)rjg+(1−pj)∑t=0∞γt(1−α)rjb =-c_j+p_j _t=0^T_p_j-1γ^t(1-α)r^g_j+(1-p_j) _t=0^∞γ^t(1-α)r^b_j =−cj+(1−α)(pjrjg1−γTpj1−γ+(1−pj)rjb1−γ) =-c_j+(1-α) (p_jr^g_j 1-γ^T_p_j1-γ+(1-p_j) r^b_j1-γ ) D(2) D^(2) =pj∑t=0Tpj−1γt+(1−pj)∑t=0∞γt=pj1−γTpj1−γ+1−pj1−γ =p_j _t=0^T_p_j-1γ^t+(1-p_j) _t=0^∞γ^t=p_j 1-γ^T_p_j1-γ+ 1-p_j1-γ Gj(2)(U;Tpj) G_j^(2)(U;T_p_j) =(1−γ)(−cj)+(1−α)(pjrjg(1−γTpj)+(1−pj)rjb)pj(1−γTpj)+(1−pj) = (1-γ)(-c_j)+(1-α) (p_jr^g_j(1-γ^T_p_j)+(1-p_j)r^b_j )p_j(1-γ^T_p_j)+(1-p_j) Under τ3 _3, we have: V(3) V^(3) =−cj+pj∑t=0Tpj−1γt(1−α)rjg+(1−pj)(1−α)rjb =-c_j+p_j _t=0^T_p_j-1γ^t(1-α)r^g_j+(1-p_j)(1-α)r^b_j =−cj+(1−α)(pjrjg1−γTpj1−γ+(1−pj)rjb) =-c_j+(1-α) (p_jr^g_j 1-γ^T_p_j1-γ+(1-p_j)r^b_j ) D(3) D^(3) =pj∑t=0Tpj−1γt+(1−pj)(γ0)=pj1−γTpj1−γ+(1−pj) =p_j _t=0^T_p_j-1γ^t+(1-p_j)(γ^0)=p_j 1-γ^T_p_j1-γ+(1-p_j) Gj(3)(U;Tpj) G_j^(3)(U;T_p_j) =−cj+(1−α)(pjrjg1−γTpj1−γ+(1−pj)rjb)pj1−γTpj1−γ+(1−pj) = -c_j+(1-α) (p_jr^g_j 1-γ^T_p_j1-γ+(1-p_j)r^b_j )p_j 1-γ^T_p_j1-γ+(1-p_j) The true Gittins index is the maximum value across all stopping rules. For our model, this is the maximum of the indices derived from these candidate rules. Gj(U;Tpj)=max0,Gj(2)(U;Tpj),Gj(3)(U;Tpj)G_j(U;T_p_j)= \0,G_j^(2)(U;T_p_j),G_j^(3)(U;T_p_j) \ Gj(G;Tpj)=rjgG_j(G;T_p_j)=r^g_j, Gj(B;Tpj)=rjbG_j(B;T_p_j)=r^b_j, and Gj(E;Tpj)=0G_j(E;T_p_j)=0 due to the persistence of the reward. B.2. Optimal Platform Policy Algorithm Algorithm 1 Optimizing platform policy _p 1: Input: Product profiles of M products 2: Generate boundary set ℬB 3: Generate regions ℛ(ℬ)R(B) 4: for each region RiR_i in ℛ(ℬ)R(B) do 5: Find Pareto optimal points in RiR_i 6: Select point _p^i that maximizes up(|)u_p( _p^i|R_i) 7: end for 8: Return policy ∗ _p^* among _p^i with highest utility B.3. Global TpT_p Toy Example Figure 2. Gittins indices for Product Type A and Product Type B as a function of global TpT_p. We consider a two-product environment where product 1 is of Type A and product 2 is of Type B, as defined in Table 1. The Gittins index of product 1 and product 2 are plotted as a function of TpT_p in Fig. 2. The zero boundary of product 1 is b10=1b_1^0=1 and the zero boundary of product 2 is b20=3.38b_2^0=3.38. The bad-indifference boundary of product 1 is b1,2B=7.23b_1,2^B=7.23 and there is no bad-indifference boundary for product 2. The undeveloped-indifference boundary b1,2U=15.27b^U_1,2=15.27. This gives us 4 regions, R1=(1,3.38)R_1=(1,3.38), R2=(3.38,7.23)R_2=(3.38,7.23), R3=(7.23,15.27R_3=(7.23,15.27), and R4=(15.27,∞)R_4=(15.27,∞). The Pareto optimal points are 1, 4, 8, and 16 for the four regions respectively. The optimal Tp∗T_p^* is Tp∗=8T_p^*=8. We also characterize the seller strategy in the following regions: • Region 1: When TpT_p is small, the seller will only explore A. The threat of immediate platform entry makes the riskier product B unattractive. • Region 2: As TpT_p increases, the seller will explore A first. If A transitions to the good state, the seller will explore B after TpT_p steps; otherwise, the seller will remain selling A in its bad state. • Region 3: As TpT_p increases even more (i.e., Tp≥8T_p≥ 8), the seller still explores A first, and the seller will choose to explore B immediately if A ends in the bad state. • Region 4: As TpT_p becomes large (i.e., Tp≥16T_p≥ 16), the seller explores Product B first. If B transitions to the good state, the seller will explore A after TpT_p; otherwise, the seller immediately switches to A. Note this is the same behavior when there is no platform entry. B.4. Heterogeneous T_p Toy Example Figure 3. Boundaries and regions induced by Product Type A and Product Type B for heterogeneous T_p We consider a two-product environment where product 1 is of Type A and product 2 is of Type B, as defined in Table 1. The boundaries and regions are labeled in Fig. 3. We do not need to optimize over the entire policy space for each RiR_i, but focus only on the Pareto optimal points where the platform’s utility upu_p is monotonic decreasing with respect to Tp1T_p_1 and Tp2T_p_2. The optimal heterogeneous entry is ∗=(1,7)T_p^*=(1,7). Appendix C Multi-seller Environment C.1. Hyperparameters IQL - DQN Param Value LR 0.0001 Discount factor 0.9 Batch Size 32 Buffer Size 450000 states Exploration Decay 0.99995 Exploration Factor 1 → 0.25 Iterative Best-Response Param Value Exploration Decay 0.999925 Exploration Factor 1 → 0.1 Total Eps 70000 ϵε (for converge) 0.33 Hyperparameters for Training: All hyperparameters omitted from Iterative Best-Response share the same hyperparameters as IQL-DQN.