Paper deep dive
AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion
Jakub SeredyÅski, Georgios Tsaousoglou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/28/2026, 4:25:26 AM
Summary
This paper investigates the emergence of tacit collusion in algorithmic electricity markets where participants use autonomous multi-agent reinforcement learning (MARL) agents for bidding. The authors hypothesize that independent learning agents can sustain supra-competitive outcomes without explicit communication or instruction. They model strategic bidding as a repeated game with imperfect public monitoring and propose a multi-dimensional set of criteria to detect collusion, including punishment of deviation, unsustainability under shortsight, and short-term profitability of deviations. Experimental results indicate that such collusion is a realistic danger in oligopolistic electricity markets.
Entities (10)
Relation Signals (7)
Multi-Agent Reinforcement Learning ā usedin ā Electricity Markets
confidence 95% Ā· We model the participantsā emergent behavior using multi-agent reinforcement learning.
Tacit Collusion ā detectedby ā Punishment of Deviation
confidence 92% Ā· This criterion checks whether an agent that deviates... will face a punitive response by other agents.
Tacit Collusion ā detectedby ā Unsustainability under Shortsight
confidence 92% Ā· The rationale of this criterion is that reducing or eliminating these elements should make it substantially more difficult for agents to learn coordinated strategies.
Tacit Collusion ā detectedby ā Short-term Profitability of Deviations
confidence 92% Ā· This criterion checks the emergent joint strategy... for these two conditions: agents having profitable unrealized deviations from it
Multi-Agent Reinforcement Learning ā leadsto ā Tacit Collusion
confidence 90% Ā· Our experimental results showcase that such a danger is realistic for electricity markets: there are cases where agents do learn to sustain supra-competitive outcomes that are supportive of tacit collusion indicators
Iterative Best Response ā usedfor ā Short-term Profitability of Deviations
confidence 88% Ā· The relevant algorithm is called āIterative Best Responseā (IBR)... to approximate Nash Equilibria
Electricity Markets ā characterizedby ā Oligopolistic Structure
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As electricity market participants increasingly adopt learning-based agents for their bidding strategies, electricity markets are becoming algorithmic. Evidence from algorithmic markets in other domains shows that tacit collusion can arise purely through independent learning. Moreover, electricity markets are typically oligopolistic and feature repeated interaction among a small number of participants, making them structurally susceptible to non-competitive behavior. In the face of these observations, this paper investigates the hypothesis that tacit collusion may emerge in electricity markets where participants' actions are controlled by autonomous learning-based algorithms. We model strategic bidding as a repeated game with imperfect public monitoring, and model the participants' emergent behavior using multi-agent reinforcement learning. We propose a multi-dimensional set of criteria (going beyond profit comparisons against Nash equilibria) to assess whether the resulting behavior constitutes tacit collusion. Our experimental results showcase that such a danger is realistic for electricity markets: there are cases where agents do learn to sustain supra-competitive outcomes that are supportive of tacit collusion indicators, even though the agents were never instructed to collude.
Tags
Links
- Source: https://arxiv.org/abs/2608.26896v1
- Canonical: https://arxiv.org/abs/2608.26896v1
Trouble viewing inline? Open PDF directly ā
Full Text
52,584 characters extracted from source content.
Expand or collapse full text
AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion Jakub SeredyÅski Georgios Tsaousoglou ā thanks: The authors are with the Department of Applied Mathematics and Computer Science, Technical University of Denmark. e-mail correspondance: geots@dtu.dk. This work was supported by the Villum Young Investigator award (Project no. 80181). Abstract As electricity market participants increasingly adopt learning-based agents for their bidding strategies, electricity markets are becoming algorithmic. Evidence from algorithmic markets in other domains shows that tacit collusion can arise purely through independent learning. Moreover, electricity markets are typically oligopolistic and feature repeated interaction among a small number of participants, making them structurally susceptible to non-competitive behavior. In the face of these observations, this paper investigates the hypothesis that tacit collusion may emerge in electricity markets where participantsā actions are controlled by autonomous learning-based algorithms. We model strategic bidding as a repeated game with imperfect public monitoring, and model the participantsā emergent behavior using multi-agent reinforcement learning. We propose a multi-dimensional set of criteria (going beyond profit comparisons against Nash equilibria) to assess whether the resulting behavior constitutes tacit collusion. Our experimental results showcase that such a danger is realistic for electricity markets: there are cases where agents do learn to sustain supra-competitive outcomes that are supportive of tacit collusion indicators, even though the agents were never instructed to collude. Index Terms: Electricity Markets, Multi-agent Reinforcement Learning, Tacit Collusion, Strategic Behavior. I Introduction I-A Motivation and Research Question Collusion in electricity markets refers to an agreement among competing firms to coordinate their actions towards avoiding competition and increasing their profits. It is severely detrimental, not only to consumers who face supra-competitive prices, but also to the marketās economic efficiency in general. While collusion is associated with cartel-like practices and is thereby strictly prohibited by law, the more challenging case to deal with is when collusion is tacit. Definition: Tacit Collusion refers to a group of oligopolistsā ability to coordinate, even in the absence of explicit agreement, in order to raise prices or, more generally, increase profits at the detriment of consumers [1]. ā” Importantly, the absence of a need for an explicit agreement makes it possible for Tacit Collusion to emerge āautonomouslyā. This is directly relevant to the new reality of electricity markets where participants are increasingly outsourcing their market participation strategies to autonomous learning-based algorithms, generally referred to as Artificial Intelligence (AI) agents. Experience from other domains has already established that AI agents, particularly Reinforcement Learning (RL), often exhibit the ability to maximize their prescribed reward (read: profit) in ways unforeseen by their designers. The study in [2] demonstrates that independently designed reinforcement-learning pricing algorithms can indeed learn to sustain supracompetitive prices, in general contexts. In the authorsā own words: āWe find that the algorithms consistently learn to charge supracompetitive prices, without communicating with one another [ā¦] The algorithms learn these strategies purely by trial and error. They are not designed or instructed to collude, they do not communicate with one another, and they have no prior knowledge of the environmentā. Taken to electricity markets, such agents could give rise to Tacit Collusion phenomena, even without the firmsā intention or awareness. Moreover, the fact that electricity markets historically tend to be oligopolistic while also featuring repeated interaction among participants, makes them particularly vulnerable to Tacit Collusion emergence. Motivated by the possibility and implications of this phenomenon in electricity markets, we put forward the following hypothesis: Hypothesis: In an electricity market where the participantsā actions are controlled by autonomous learning-based algorithms, Tacit Collusion may emerge. I-B Related Work An electricity market comprises participants, each of which is interested in optimizing its own profit. Thereby, studying possibilities over electricity market outcomes involves modeling the bidding strategy optimization problem of individual participants, in an agent-based modeling fashion [3]. In the earlier literature, a participantās bidding strategy was predominantly modeled through bi-level programming (e.g. [4]), giving rise to mixed-complementarity problems modeling the multi-participant interactive decision-making [5, 6]. The recent literature, however, has substantially pivoted toward learning-based approaches, for at least three reasons: ⢠Learning algorithms are able to handle realistic aspects of electricity markets such as complexity, partial observability, and stochasticity, departing from simplifying assumptions arising from computational constraints of the model [7]. ⢠The model-free nature of learning algorithms makes them a generic modeling tool which, taken to a sufficient level of richness, can in principle reproduce a wide range of different intelligent strategies [8, 9]. ⢠Real market participants increasingly experiment with learning-based bidding strategies [10]. This motivates the RL framework, in particular, to be leveraged as a suitable modeling framework that captures the exploration aspect of modern algorithmic trading. RL has attracted special interest as a generic model for a participantās bidding strategy optimization problem [11], [12], [13]. From a system perspective, when it comes to modeling the interactive learning among several strategic participants and the emergent outcomes thereof, the concept naturally evolves into multi-agent RL (MARL) [14]. To that end, [15] argues for MARL-based fit-for-purpose simulators that retain the market and physical features relevant to the research question at hand, while [16] formulates day-ahead bidding with multiple strategic generators as algorithmic traders and applies a MARL algorithm to approximate Nash-equilibrium bidding. Recent work on generator modeling further shows that simplified bidding and marginal-cost representations can affect conclusions in market-mechanism simulations, especially when renewable integration changes the operating regimes of thermal generators [17]. The potential emergence of Tacit Collusion has been validated by research in other domains, namely algorithmic pricing, where RL agents have been shown to sustain supra-competitive prices without communication or explicit collusive instructions [2]. This result has been extended to sequential pricing environments, where Q-learning algorithms may converge either to collusive equilibria or to supra-competitive price cycles [18]. Recent experimental evidence further compares algorithmic and human collusion in the same market environments, showing that self-learning pricing algorithms can generate more collusive outcomes than human decision-makers in some oligopoly settings [19]. Taken to the electricity-market setting, it is particularly timely to understand the implications of algorithmic trading and, in particular, whether Tacit Collusion can emerge endogenously from the interaction of autonomous learning agents. To that end, a MARL agent-based model is employed in [3], where the authors investigate how the agentsā discount factors affect the resulting market outcomes. More recently, [20] explicitly investigates the emergence of Tacit Collusion and define a collusive state as one in which the profit of each participant exceeds that attained under all Nash equilibria. While this provides a useful benchmark for identifying supra-competitive outcomes, establishing such a condition requires the computation of the Nash equilibria of the underlying market game. This requirement substantially constrains the complexity of the market model and makes the approach difficult to extend to richer electricity-market settings. I-C Research Gap and Contributions The emergence of algorithmic trading, and the risks it poses for electricity markets, points to a major research need around the emergence of Tacit Collusion. To our knowledge, this phenomenon has been addressed by only a handful of studies (namely [20] and [3]). These studies provide important evidence that intelligent agents may learn supra-competitive bidding strategies. Their identification of Tacit Collusion, however, is closely tied to the structure of the underlying market game. In particular, [20] identifies collusion by comparing the agentsā profits with those attained at all Nash equilibria. Such an approach is informative in a stylized market model whose equilibria can be fully characterized, but becomes difficult to apply once the market model incorporates the structural features of actual electricity markets. More fundamentally, supra-competitive profits alone do not establish Tacit Collusion: elevated profits can arise for reasons unrelated to collusive coordination. Conversely, because realistic market models rarely converge to a stage-game Nash equilibrium, the absence of such convergence is, on its own, equally uninformative about whether collusion has occurred. In fact, in our numerical results in this paper, we present examples on both ends: cases where supra-competitive outcomes emerge which are, or are not necessarily, supportive of a case for Tacit Collusion. Tacit Collusion is thus a subtle phenomenon that resists identification through any single criterion and calls instead for a multi-dimensional examination. This motivates the contributions of this paper, positioned towards investigating the Hypothesis defined in Section I-A: 1. We adopt Tacit Collusion criteria (beyond merely supra-competitive outcomes) from other domains, with richer experience in this context, and adapt them to the electricity market setting. We also propose a novel criterion which quantifies the central characteristic of collusion as defined in its Definition (cf Section I-A), in the context of electricity markets. 2. We model emergent bidding profiles using a MARL framework and examine their collusive properties across multiple such criteria. 3. We experimentally investigate which market characteristics (e.g. grid constraintsā tightness, demand level, marginal costsā variance) enhance the possibility of emergent collusion. The remainder of the paper is organized as follows. Section I presents the system model. Section I defines the different Tacit Collusion criteria of our study. Section IV presents the experimental setting, including the details of the MARL algorithm. Section V presents our experimental results, while Section VI concludes the paper. I System Model I-A Electricity Market We consider a single-timeslot electricity market, operating over a set N of buses, which is cleared by solving the standard DC optimal power flow (OPF) problem. Specifically, a bus nān features an inelastic demand DnD_n and a set ānI_n of energy-providing resources. The resources can generally be of various types and technologies (including conventional power plants, renewable energy sources, storage, flexible demand, etc) but we will simply call them āgeneratorsā to simplify the exposition. We denote the superset of generators (across all nodes) by ā=ānāānI= _n I_n. The active power flow from bus n to bus m is denoted as fnāmf_nm and is subject to line capacity bounds per āFnāmā¤fnāmā¤Fnām,ā(n,m)ā2,-F_nm⤠f_nm _nm, ā(n,m) ^2, (1) where, for pairs of buses that are not connected by a line, we simply set Fnām=0F_nm=0. Denoting the output of a generator iāāni _n as qiq_i, the power balance constraint at a node reads: āiāānqiāāmāfnām=Dn,ānā. _i _nq_i- _m f_nm=D_n, ā n . (2) The phase angle at a bus n is denoted as Ļn _n and is subject to safety limits per ĻĀÆnā¤Ļnā¤ĻĀÆn,ānā. Ļ_n⤠_n⤠Ļ_n, ā n . (3) The phase angles and the power flows are linked by the standard DC power flow model: fnām=Bnāmā(ĻnāĻm),ā(n,m)ā2,f_nm=B_nm( _n- _m), ā(n,m) ^2, (4) where BnāmB_nm is the susceptance of transmission line nāmnm. The output of a generator iāāi is partitioned into a set iS_i of segments, as in: qi=āsāiqi,s,āiāā,q_i= _s _iq_i,s, ā i , (5) where the output qi,sq_i,s of a segment is bounded by 0ā¤qi,sā¤Qs,āsāi,iāā.0⤠q_i,s _s, ā s _i,i . (6) Finally, each generator declares a per-unit energy cost ci,sc_i,s for each of its segments sāis _i, such that the operator receives a piecewise linear cost function for each generator. The market is cleared by solving the standard DC-OPF problem: minqi,qi,s,fnām,Ļnāiāāāsāi q_i,q_i,s,f_nm, _nmin~ \ _i _s _i ci,sqi,s _i,sq_i,s \ (7) subject to (1)ā(4) c-flows- c-dc :Network constraints :Network constraints (5)ā(6) c-segsum- c-seg :Generatorsā constraints. :Generators' constraints. The optimal dual variable Ī»n _n of the power balance constraint (2) instantiates the Locational Marginal Price (LMP) of node n. The market-clearing problem (7) is implemented sequentially for a set of instances (timeslots) T where in each instance tāt the generators can submit different bids ci,sā[t]c_i,s[t] resulting in timeslot-specific solutions ((qiāā[t])iāā,(Ī»nā[t])nā) ((q_i^*[t])_i ,( _n[t])_n ). I-B Agent-based Market Participation Strategies A generatorās market participation strategy is instantiated by an agent. We use the terms āagentā and āgeneratorā interchangeably and we index either by i. After each market-clearing instance t, the stage profit Ļiā[t] _i[t] of i is Ļiā[t]=Ī»nā”(i)ā[t]āqiāā[t]āCiā(qiāā[t]). _i[t]= _n(i)[t]q_i^*[t]-C_i(q_i^*[t]). (8) The first term is the generatorās revenue, with Ī»nā”(i)ā[t] _n(i)[t] being the LMP at bus nā”(i)n(i) where i is located. The second term is the generatorās real cost for generating its dispatched quantity qiāā[t]q_i^*[t], which is given by the generatorās cost function Ci:āāāC_i:R . The values for qiāā[t]q_i^*[t] and Ī»nā”(i)ā[t] _n(i)[t] are given from problem (7), and are therefore affected by the actions, i.e. bids, (ci,sā[t])sāi(c_i,s[t])_s _i of the focal agent i, as well as by those of the other agents. This directly leads to the definition of the stage game G: ⢠Players: iāāi ⢠Actions: iā[t]=(ci,sā[t])sāi c_i[t]=(c_i,s[t])_s _i ⢠Payoffs: Ļiā[t] _i[t] as in Eq. (8) We consider agents that learn to improve their strategies through experience. After each instance, an agent gets to observe its own dispatch and profit, as well as the resulting LMP at every node of the network: oiā[t]=(qiāā[t],Ļiā[t],(Ī»nā[t])nā).o_i[t]= (q_i^*[t], _i[t],( _n[t])_n ). (9) This gives rise to a repeated game with imperfect monitoring, since an agent does not observe the actions or profits of other agents but only a public signal (i.e. the LMPs) from which the othersā actions and profits cannot be inferred. Furthermore, we do not assume that agents know the structure of the network or even the market mechanism and consider agents that learn purely through experience. To that end, each agent accumulates a private history of its actions and observations (up to ζ episodes behind): hiā[t]=(iā[tā1],oiā[tā1],iā[tā2],oiā[tā2],ā¦,iā[tāζ],oiā[tāζ])h_i[t]= ( c_i[t-1],o_i[t-1], c_i[t-2],o_i[t-2],..., c_i[t-ζ],o_i[t-ζ] ) (10) and selects its bid at each stage t according to a strategy Ļi _i that maps the current history hiā[t]h_i[t] to its decision iā[t] c_i[t]. We denote the joint strategy as =(Ļi)iāā Ļ=( _i)_i . Based on the above, we can define the objective of an agent as a search for a strategy that maximizes its cumulative profits over time: maxĻiāāt=0āĻiā[t] _imax~ \E_ Ļ _t=0^ā _i[t] \ (11) subject to (8):Profits depend on actions, profit:Profits depend on actions, iā[t]=Ļiā(hiā[t]):Actions are given by strategy and history c_i[t]= _i(h_i[t]):Actions are given by strategy and history (10):The history stacks part actions and observations history:The history stacks part actions and observations (9):Observations are: dispatch, profit, and LMPs. observation:Observations are: dispatch, profit, and LMPs. Problem (11) is intractable in general: each agent optimizes its own strategy while the strategies of all other agents are unknown and simultaneously evolving, making the environment non-stationary from any single agentās perspective. In line with the learning-through-experience consideration, we adopt Multi-Agent Reinforcement Learning (MARL) as a model of how AI agents learn and play in practice. At each instance t, every agent i selects a bid iā[t]=Ļiā(hiā[t]) c_i[t]= _i(h_i[t]), the market clears, and each agent receives the observation oiā[t]o_i[t] and uses it to update its strategy Ļi _i. This process is iterated over successive instances, and the agentsā strategies co-evolve. The details of the MARL algorithm will be presented in Section IV-C. The output of this process is a joint strategy profile ā=(Ļiā)iāā Ļ^*=( _i^*)_i , representing the emergent bidding behaviour that arises when self-interested agents learn to bid. The paperās Hypothesis thereby refers to analyzing the emergent joint strategies ā Ļ^* in terms of indications for Tacit Collusion. The next section formalizes the Tacit Collusion indices to be used. I Tacit Collusion In this section, we define three criteria that constitute indicators of the emergent joint strategy ā Ļ^* being tacitly collusive: Punishment of Deviation, Unsustainability under Shortsight, and Short-term Profitability of Deviations. In the three subsection below, we describe how each criterion is instantiated in our electricity market context. I-A Punishment of Deviation Criterion This criterion checks whether an agent that deviates from the joint strategy ā Ļ^* and reverts to an explicitly competitive strategy will face a punitive response by other agents. A deviator reverting to a competitive strategy means that it starts to optimize its short-term profits aggressively. In our context, we can model such a situation by choosing one agent j as the deviator and, after the MARL strategies have converged to ā Ļ^*, switching the deviatorās strategy to the one that optimizes its stage profit under the assumption that other agents iā jiā j continue to play their previous strategies (specifically, their most recently observed ones). To implement this, at each after-convergence stage t, we force jās strategy to the solution of the following bi-level optimization problem: maxjā[t]āĪ»nā”(j)ā[t]āqjāā[t]āCjā(qjāā[t]) c_j[t]max~ \ _n(j)[t]q^*_j[t]-C_j(q^*_j[t]) \ (12) subject to qjāā[t],Ī»nā”(j)ā(7) q^*_j[t], _n(j)ā market-clearing where the constraint specifies that the dispatch and LMP result through the market clearing problem (7). Problem (12) can be solved by casting the lower-level problem (7) as a set of KKT conditions and introducing auxiliary binary variables to handle the resulting complementarity constraints. This results in a single-level mixed-integer linear program which can be tackled by commercial solvers. The full derivation is ommited here as it is fairly standard in the literature. While the deviator is using its aggressive profit-maximizing strategy (12), the other agents continue to play by the MARL regime. A punitive response is considered to be one that has other agents lowering their mark-ups in a way that reduces the deviatorās profit. The test is considered positive if the post-deviation dynamics display the qualitative pattern of punish-and-forgive strategies: ⢠Immediately after the deviation, competitors temporarily switch to more competitive (i.e. lower) bids, reducing prices and the deviatorās market share; ⢠After a punishment phase, competitors gradually return to their pre-deviation strategies, while prices recover. I-B Unsustainability under Shortsight Criterion Two elements that generally facilitate Tacit Collusion are: ⢠agents playing the long game, by adopting a high discount factor which increases the importance of future payoffs over immediate ones; ⢠agents having memory of past actions and rewards, which allows them to condition their behaviour on history; The rationale of this criterion is that reducing or eliminating these elements should make it substantially more difficult for agents to learn coordinated strategies. In practice, this test is implemented by executing the MARL algorithm under a very low discount factor while also severely truncating the memory of past actions and rewards. Compared to the baseline case, this run would presumably result in more competitive strategies, i.e. lower prices and profits. I-C Short-term Profitability of Deviations Criterion Given that collusive outcomes are not Nash equilibria of the stage game, one or more agents do have a profitable unilateral strategy deviation. Yet the collusion is sustained when agents refrain from switching to opportunistic strategies, despite this being profitable in the short-run, presumably because they understand that those would lead to down-the-road less profitable outcomes. This criterion checks the emergent joint strategy ā Ļ^* precisely for these two conditions: agents having profitable unrealized deviations from it, and whether those, if realized, would trigger joint adaptation dynamics that indeed lead to less profitable trajectories. Similarly to the first criterion, we instantiate profitable deviations by solving the bi-level problem (12). This time, however, we solve problem (12) iteratively and for every agent in turn, thereby simulating a relapse into competitive (non-collusive) strategy adaptation. The relevant algorithm is called āIterative Best Responseā (IBR) or ādiagonalizationā and it is widely used in the electricity marketsā literature (see e.g. [21] and references therein) to approximate Nash Equilibria under the so-called Equilibrium-Problem with Equilibrium-Constraints (EPEC) model. The criterion considers that ā Ļ^* passes the collusion test if, upon switching from MARL to the IBR regime, the profits of the first deviating agents: ⢠increase at first, and ⢠they eventually drop below those under ā Ļ^*, as the IBR process continues towards a Nash Equilbirium. The rationale is that, if both of these conditions hold true, it means that each agent has a profitable unilateral deviation which it doesnāt realize, possibly because it foresees that it is not profitable under the other agentsā corresponding competitive response, thereby staying true to a coordinated (collusive) strategy. IV Experimental Setting IV-A Electricity Market Setup We conduct the experiments on a stylised seven-bus electricity market adapted from the case study of [20]. The original system is an extended version of the well-known PennsylvaniaāNew-JerseyāMaryland (PJM) five-node test system and was designed as a compact networked market in which strong collusive equilibria can arise. We use it here as a controlled benchmark that preserves the elements needed for our hypothesis: an oligopolistic supply side, inelastic nodal demand, transmission constraints, and locational marginal prices. The market contains four strategic generators located at buses N1N_1, N2N_2, N5N_5, and N6N_6. Nodal demand is perfectly inelastic and the market is cleared, at each instance, by the DC-OPF problem (7). The resulting dispatch and nodal prices determine each generatorās stage profit according to (8), where revenues are computed at the LMP of the generatorās bus and costs are computed using the generatorās true marginal cost. The main generation and cost parameters used in the experiments are reported in Table I. TABLE I: Strategic generators and marginal-cost cases. Generator Bus PimaxP_i [MW] Marginal cost [EUR/MWh] Low Medium High G1G_1 N1N_1 42 20 20 20 G2G_2 N2N_2 35 20 20 20 G5G_5 N5N_5 40 20 30 40 G6G_6 N6N_6 40 20 10 1 All exogenous fundamentals are kept fixed within a scenario. In particular, nodal demand, marginal costs, and network parameters do not vary over time. The three demand cases, indexed by buses (N1,ā¦,N7)(N_1,ā¦,N_7), are DL=(17,10,8,6,13,0,12)D^L=(17,10,8,6,13,0,12), DM=(22,14,10,8,18,0,13)D^M=(22,14,10,8,18,0,13), and DH=(26,15,12,9,20,0,14)D^H=(26,15,12,9,20,0,14). Thus, the temporal dynamics observed in the simulations are induced by the repeated interaction of bidding strategies and by the feedback of the market-clearing mechanism, rather than by exogenous factors. To isolate the role of transmission constraints, we consider two network representations. In the Grid case, the network constraints (1)ā(4) are enforced using the topology and line limits of the test system. In the NoGrid case, the same demand and generator data are retained, but line capacities are set sufficiently high so that the bounds in (1) never become binding, while the corresponding admittance values are chosen so that the DC flow relation (4) and phase-angle bounds (3) do not restrict the optimal dispatch. Operationally, this produces a copper-plate counterfactual. Mathematically, for all solved instances, the NoGrid case is equivalent to relaxing the network-induced restrictions in (1), (3), and (4), so that problem (7) clears the market as a single unconstrained zone. The experimental design is a full factorial combination of three axes: network representation Grid/NoGrid Grid/ NoGrid, demand level Low/Medium/High Low/ Medium/ High, and marginal-cost heterogeneity Low/Medium/High Low/ Medium/ High. This gives 2Ć3Ć3=182Ć 3Ć 3=18 market environments. The selected scenario dimensions correspond to market characteristics that are widely discussed in the Tacit Collusion literature as factors influencing the emergence and sustainability of collusive outcomes, namely network topology, demand conditions, and cost heterogeneity [1, 22]. All scenarios are trained and screened, while the behavioural collusion tests are applied only to the most informative cases selected according to the procedure described next. IV-B Scenario Selection Methodology Instead of running our tests in randomly selected market instances, we identify instances that are most suspicious for potential Tacit Collusion emergence. In this subsection, we describe the methodology that we use to select which market instances will be tested for the criteria of Section I. Our starting point is the 18 scenarios of the full factorial design described in the previous subsection. For each of them, we compute screening indicators over the final evaluation window of each training run, after exploratory noise has decayed. The competitive benchmark is obtained by solving the same market-clearing problem (7) under marginal-cost bidding, i.e. with all generators bidding their true marginal costs. Denoting the average realised price and total profit by λ¯oābās Ī»^obs and Ī oābās ^obs, and their competitive counterparts by λ¯cāoāmāp Ī»^comp and Ī cāoāmāp ^comp, we define MĪ»=λ¯oābāsāλ¯cāoāmāpλ¯cāoāmāp,M^Ī»= Ī»^obs- Ī»^comp Ī»^comp, (13) MĪ =Ī oābāsāĪ cāoāmāpmaxā”Ī cāoāmāp,ϵΠ,M = ^obs- ^comp \ ^comp, _ \, (14) where Ī =āiāIĻi = _iā I _i denotes aggregate generator profit and ϵΠ_ is a small denominator floor used to avoid degenerate ratios when competitive profits are close to zero. To separate high-price outcomes from coordinated behaviour, the markup indicators are complemented by three dynamic indicators. First, price variance measures whether the learned market regime is stable enough for an intervention test to be meaningful. Second, action correlation and action-change correlation measure, respectively, whether agents maintain similar bidding levels or move their bidding strategies together over time. Third, Best Response Deviation Gain (BRDG) measures the short-run profitability of unilateral deviations. For agent i, it is defined as BRDGi=Ļiā(ciBāR,cĀÆāi)āĻiā(cĀÆi,cĀÆāi)Ļiā(cĀÆi,cĀÆāi),BRDG_i= _i(c_i^BR, c_-i)- _i( c_i, c_-i) _i( c_i, c_-i), (15) where cĀÆi c_i denotes the bid prescribed by the learned policy of agent i, cĀÆāi c_-i denotes the learned bids of all other agents, and ciBāRc_i^BR is the one-shot best response to cĀÆāi c_-i. A positive BRDG therefore indicates that the learned profile contains an unrealised short-run deviation incentive. The selection favours cases that satisfy three practical requirements: (i) the learned regime is sufficiently stable to be meaningfully perturbed, (i) the outcome is not a degenerate boundary solution in which all agents simply bid at the action-space limit, and (i) there is a non-trivial unilateral deviation incentive that can be tested through the criteria of Section I. The selected cases are then evaluated using the Punishment of Deviation Criterion, the Unsustainability under Shortsight Criterion, and the Short-term Profitability of Deviations Criterion described in Section I. TABLE I: Scenarios selected for behavioural validation. Scenario Demand Cost diff. Rationale GridāLowDemandāHighCostDiff Low High Structured non-stationary regime with non-trivial deviation incentives. GridāLowDemandāMediumCostDiff Low Medium Reference low-demand case with a more stationary learned profile. GridāMediumDemandāHighCostDiff Medium High Higher-demand case that remains non-boundary and strategically informative. IV-C MARL implementation The action of agent i at each instance t is to select a bid vector iā[t]=(ci,sā[t])sāi c_i[t]=(c_i,s[t])_s _i, i.e. one declared marginal cost per segment s. This renders the action space of each agent to be of dimension |i||S_i|. To reduce this dimensionality and facilitate learning, we parameterize the bid vector through a supply function. Specifically, each agent i maintains a linear supply function fā”(qi,βi)=αi+βiā qiāsāiQs,f(q_i; _i)= _i+ _iĀ· q_i _s _iQ_s, (16) which maps an output level qi,sq_i,s to a declared marginal cost (bid) ci,sc_i,s. Note that this function is generally different than the agentās true cost function (defined in the previous subsection); it is an auxiliary bidding curve used internally to generate a consistent set of segment bids. The bid ci,sc_i,s for segment s is obtained by evaluating f at the midpoint qĀÆi,s=āsā²ā0,1,ā¦āsā1Qsā²+Qs2 q_i,s= _s ā\0,1,...s-1\Q_s + Q_s2 of the segmentās capacity range: ci,sā[t]=fiā(qĀÆi,s,βi)=αi+βiā qĀÆi,sāsāiQs,sāi,c_i,s[t]=f_i( q_i,s; _i)= _i+ _iĀ· q_i,s _s _iQ_s, s _i, (17) where αi _i is set to the agentās true marginal cost and βi _i is the parameter to be learned by the agent. This defines a deterministic mapping βiā[t]ā¦iā[t] _i[t] c_i[t], so that the full bid vector is determined by a single scalar parameter. The agentās internal action at each instance t is therefore the choice of βiā[t]āāā„0 _i[t] _ā„ 0, which is then translated into the game action iā[t] c_i[t] via (17). Remark 1 This parameterization reduces the dimensionality of an agentās action space from |i||S_i| to 11. This is a choice that only makes the validation of our Hypothesis more difficult: by reducing the agentsā degrees of freedom, we reduce the richness of their policies, possibly excluding instances of collusive joint policies. Thus, if collusive policies are not discovered with this model, it does not necessarily mean that they donāt exist. But if collusive policies are discovered with this model, then the validation of the Hypothesis holds also for the general model where each agentās action comprises an arbitrary choice for each segmentās bid. ā” In the experiments, the supply function of each generator is discretised into 12 equal-capacity bid segments. The bidding parameter βi _i is constrained to the interval [0,4000][0,4000], where the upper bound corresponds to the maximum admissible value of the slope parameter. The upper bound is consistent with the market price cap of 4000 EUR/MWh. The agents are trained using a multi-agent implementation of TD3 under the centralized-training/decentralized-execution paradigm. Each generator is represented by a deterministic actor that maps its private history to the bidding parameter used to construct its offer. During training, critics are allowed to condition on joint market information in order to stabilize learning in the non-stationary multi-agent environment. During execution and evaluation, however, only the decentralized actors are used. The agents therefore do not communicate when submitting bids. Each of the 18 scenarios is trained for 100,000 market instances, organised as 50 episodes of 2,000 instances. The first 4,000 instances are used as a warm-up phase, during which agents execute randomly generated actions to populate the replay buffer. Exploration is introduced through action noise and is gradually removed over the course of training, so that the final evaluation window reflects the learned deterministic policies rather than exploratory behaviour. All screening indicators reported in the previous subsection are computed over the final 10,000 market instances. The main hyperparameters are reported in Table I. TABLE I: Main MARL hyperparameters. Parameter Value Algorithm Multi-agent TD3 Bid segments 12 Bidding parameter bounds [0,4000][0,4000] Training length 100,000 market instances Episodes / horizon 50Ć2,00050Ć 2,000 Evaluation window Final 10,000 instances Warm-up phase 4,000 instances Discount factor 0.99 Replay buffer size 50,000 Batch size 256 Actor / critic learning rate 10ā410^-4 / 3ā 10ā43Ā· 10^-4 Hidden layers [256,256][256,256] Exploration decay 90,000 instances V Results V-A Screening Outcomes Among the screening indicators introduced in Section IV, Fig. 1 reports the price markup across the full set of 18 market environments. Markups are generally modest in the NoGrid cases and do not exhibit a clear ordering across demand levels or cost heterogeneity. In contrast, the Grid cases produce substantially higher supra-competitive markups and display a systematic pattern in which higher demand and lower marginal-cost heterogeneity are associated with higher markups. This indicates that transmission constraints materially amplify market power and make the resulting market outcomes more sensitive to the underlying demand and cost structure. At the same time, elevated markups alone cannot distinguish coordinated behaviour from market power arising directly from network constraints. (a) (b) Fig. 1: Price markups across the nine demand and marginal-cost heterogeneity combinations: (a) NoGrid scenarios without active transmission constraints and (b) Grid scenarios with active transmission constraints. Applying the screening procedure described in Section IV leaves three non-boundary Grid cases for behavioural validation: Grid-LowDemand-HighCostDiff, Grid-LowDemand-MediumCostDiff, and Grid-MediumDemand-HighCostDiff. These cases all exhibit supra-competitive outcomes while differing in demand conditions, cost asymmetry, and learned strategic dynamics. V-B Punishment of Deviations Figure 2 shows the bidding responses to the forced deviation in the three selected scenarios. In Grid-LowDemand-HighCostDiff, the deviation of G1G_1 triggers a pronounced response from the remaining agents, which temporarily lower their bidding slopes and subsequently return towards their pre-deviation strategies, producing a clear punish-and-forgive pattern. A similar pattern emerges in Grid-LowDemand-MediumCostDiff, although the response is asymmetric: G5G_5 and G6G_6 react strongly, while G2G_2 remains close to its pre-deviation strategy. A plausible explanation for this difference is the change in the residual-supply structure induced by the different relative competitiveness of G5G_5 and G6G_6, rather than any change in G2G_2ās own marginal cost, which is identical in the two scenarios. Under HighCostDiff, the very competitive G6G_6 is strongly constrained by its network position while G5G_5 is a relatively costly substitute, increasing the marginal value of G2G_2ās available capacity and making its participation in the punishment more important. Under MediumCostDiff, G5G_5 becomes more competitive and G6G_6 retains more marginal supply headroom, so their combined response can substitute more effectively for G2G_2, reducing the marginal value of its participation and creating a natural free-riding incentive. In Grid-MediumDemand-HighCostDiff, by contrast, the deviation of G5G_5 does not induce a comparable coordinated reduction in competitorsā bidding slopes, and no clear punish-and-forgive pattern is observed. Fig. 2: Agentsā bidding actions during the Punishment of Deviation test for the three selected scenarios. The shaded region denotes the forced-deviation window, and the black line identifies the deviating agent. The corresponding reward trajectories in Fig. 3 show whether these strategic responses impose an economic penalty on the deviating agent. In Grid-LowDemand-HighCostDiff, the response of the other agents reduces the profitability of G1G_1ās deviation, with rewards subsequently recovering as the agents return towards the previous regime. In Grid-LowDemand-MediumCostDiff, the response of G5G_5 and G6G_6 is likewise sufficient to penalize the deviator despite the weak reaction of G2G_2, supporting the interpretation that G2G_2 is not pivotal to punishment in this case. In Grid-MediumDemand-HighCostDiff, the absence of a comparable strategic response is reflected in the deviator retaining the benefit of its deviation, providing no evidence of an effective punishment mechanism. The Punishment of Deviation criterion is therefore supported in both LowDemand scenarios, although with an asymmetric response in the MediumCostDiff case, and is not supported in Grid-MediumDemand-HighCostDiff. Fig. 3: Agentsā rewards during the Punishment of Deviation test for the three selected scenarios. The shaded region denotes the forced-deviation window, and the black line identifies the deviating agent. V-C Shortsight Ablation Figure 4 compares the evolution of episode-average bidding slopes under the baseline MARL configuration and the shortsight ablation. Contrary to a simple collapse-to-competition interpretation, removing memory and long-horizon incentives does not systematically reduce the learned bidding slopes: the magnitude and direction of the changes vary across agents and scenarios, while the Grid-MediumDemand-HighCostDiff case remains particularly close to the baseline. A plausible explanation is that a substantial part of the on-path supra-competitive outcome is supported by static, network-induced market power. Transmission constraints, local scarcity, and generator pivotality remain unchanged after the ablation, so high bidding slopes may remain profitable even for nearly myopic agents. The stationarity of demand, costs, and network parameters further allows the learned policies to stabilize around similar actions without relying on long histories of past market outcomes. The ablated runs also converge more smoothly, consistent with weaker inter-agent feedback loops. Hence, the similarity of the on-path outcomes does not imply that the underlying strategic mechanism is unchanged; the relevant distinction is between the level of the learned regime and the agentsā ability to condition their behaviour on deviations from it. Fig. 4: Evolution of episode-average bidding slopes under the baseline MARL configuration and the shortsight ablation for the three selected scenarios. Solid lines denote the baseline agents and dashed lines the ablated agents. This distinction becomes apparent when the forced-deviation experiment is repeated under the shortsight ablation, as shown in Fig. 5. In both LowDemand scenarios, the pronounced punishment-and-forgiveness responses observed under the baseline disappear, with competitorsā bidding slopes remaining largely unchanged following the forced deviation. Removing history eliminates the informational basis for conditioning current actions on a previous deviation, while the near-myopic discount factor removes much of the incentive to incur a short-run punishment cost in order to restore a more profitable future regime. In Grid-MediumDemand-HighCostDiff, where no clear punishment response was present under the baseline, the ablation likewise produces no evidence of an enforcement mechanism. The shortsight ablation therefore affects primarily the off-path enforcement mechanism rather than necessarily the on-path level of bids. It consequently provides supporting evidence for a repeated-game component in the two LowDemand scenarios through the disappearance of punishment, while the persistence of similar on-path bidding levels indicates that their supra-competitive outcomes are also partly supported by static market power; no comparable evidence is obtained for Grid-MediumDemand-HighCostDiff. Fig. 5: Agentsā bidding actions during the Punishment of Deviation test under the shortsight ablation for the three selected scenarios. The shaded region denotes the forced-deviation window, and the black line identifies the deviating agent. V-D Iterative Best Responses Figure 6 shows the reward trajectories obtained when the learned MARL outcomes are used as the starting point for Iterative Best Response. In Grid-LowDemand-HighCostDiff, unilateral best responses provide short-run profit opportunities, but the subsequent IBR dynamics move the agents towards reward levels that are generally below those sustained under MARL. The pattern is even clearer in Grid-LowDemand-MediumCostDiff, where the initially profitable deviations are followed by lower reward levels for all agents relative to their MARL benchmarks. These two cases are therefore consistent with a regime in which unilateral deviation is attractive in the short run, while the competitive adaptation induced by continued best responses is less profitable. In Grid-MediumDemand-HighCostDiff, by contrast, the IBR trajectories do not exhibit the same combination of profitable short-run deviations and a systematically less profitable subsequent regime. The Short-term Profitability of Deviations criterion is therefore supported in the two LowDemand scenarios and not supported in Grid-MediumDemand-HighCostDiff. Fig. 6: Agent reward trajectories under Iterative Best Response (IBR), initialized from the learned MARL outcomes. Iteration 0 corresponds to the MARL endpoint, while the horizontal dashed lines indicate the corresponding MARL reward levels. V-E Synthesis of Behavioural Evidence TABLE IV: Summary of behavioural evidence for Tacit Collusion. Scenario Punish. Shortsight IBR Interpretation GridāLowDemandāHighCostDiff Supportive Supportive Supportive Collusion-consistent GridāLowDemandāMediumCostDiff Supportive Supportive Supportive Collusion-consistent GridāMediumDemandāHighCostDiff Not supp. Not supp. Not supp. Structural market power Table IV summarizes the evidence obtained from the three behavioural criteria. Taken together, the three behavioural tests provide mutually consistent evidence that the two LowDemand learned regimes contain a repeated-game component consistent with Tacit Collusion. The persistence of supra-competitive on-path bidding after the shortsight ablation nevertheless indicates that these outcomes are not driven by repeated-game incentives alone, but are also supported by structural, network-induced market power. Grid-MediumDemand-HighCostDiff provides the contrasting case: despite exhibiting a supra-competitive market outcome, it shows neither a punishment response, an ablation-sensitive enforcement mechanism, nor an IBR pattern consistent with the collusion criteria. Its elevated markup is therefore more plausibly attributed to structural market power than to Tacit Collusion. Overall, the results show that similar supra-competitive outcomes can arise from qualitatively different strategic mechanisms, which cannot be distinguished from price levels alone. VI Conclusions This paper investigated whether Tacit Collusion can emerge in electricity markets whose participants delegate their bidding strategies to autonomous, learning-based algorithms. Strategic bidding was modeled as a repeated game with imperfect public monitoring, with agentsā learning-based policies modeled via multi-agent reinforcement learning. Different cases of market characteristics were filtered based on susceptibility for Tacit Collusion and the most suspicious ones were tested against three Tacit Collusion criteria. Our experimental results show that the danger of Tacit Collusion in electricity markets is realistic: particularly under binding network constraints, agents learn to sustain supra-competitive outcomes that satisfy several independent indicators of Tacit Collusion, despite never being instructed or designed to collude. Our results should not be taken as conclusive but rather as an alarm bell that encourages extensive research on this topic. Specifically, investigating the learning mechanics that make Tacit Collusion emergent and investigating market-design countermeasures constitute important future work directions. References [1] M. Ivaldi, B. Jullien, P. Rey, P. Seabright, and J. Tirole (2007) The economics of tacit collusion: implications for merger control. In The Political Economy of Antitrust, p. 217ā239. External Links: ISBN 9780444530936, ISSN 0573-8555, Link, Document Cited by: §I-A, §IV-A. [2] E. Calvano, G. Calzolari, V. Denicolò, and S. Pastorello (2020) Artificial intelligence, algorithmic pricing, and collusion. American Economic Review 110 (10), p. 3267ā3297. External Links: ISSN 0002-8282, Link, Document Cited by: §I-A, §I-B. [3] Y. Liang, C. Guo, Z. Ding, and H. Hua (2020) Agent-based modeling in electricity market using deep deterministic policy gradient algorithm. IEEE transactions on power systems 35 (6), p. 4180ā4192. Cited by: §I-B, §I-B, §I-C. [4] E. G. Kardakos, C. K. Simoglou, and A. G. Bakirtzis (2014) Optimal bidding strategy in transmission-constrained electricity markets. Electric Power Systems Research 109, p. 141ā149. Cited by: §I-B. [5] X. Hu and D. Ralph (2007) Using epecs to model bilevel games in restructured electricity markets with locational prices. Operations research 55 (5), p. 809ā827. Cited by: §I-B. [6] S. A. Gabriel, A. J. Conejo, J. D. Fuller, and B. F. Hobbs (2013) Complementarity modeling in energy markets. Springer New York, NY. Cited by: §I-B. [7] G. Tsaousoglou, J. S. Giraldo, and N. G. Paterakis (2022) Market mechanisms for local electricity markets: a review of models, solution concepts and algorithmic techniques. Renewable and Sustainable Energy Reviews 156, p. 111890. Cited by: 1st item. [8] Q. Tang, H. Guo, K. Zheng, and Q. Chen (2024) Forecasting individual bids in real electricity markets through machine learning framework. Applied Energy 363, p. 123053. Cited by: 2nd item. [9] S. Baltaoglu, L. Tong, and Q. Zhao (2018) Algorithmic bidding for virtual trading in electricity markets. IEEE Transactions on Power Systems 34 (1), p. 535ā543. Cited by: 2nd item. [10] N. Eschenbaum (2026) Shared bidding algorithms and competition: evidence from electricity markets. External Links: 2607.13002, Link Cited by: 3rd item. [11] Y. Ye, D. Qiu, M. Sun, D. Papadaskalopoulos, and G. Strbac (2019) Deep reinforcement learning for strategic bidding in electricity markets. IEEE Transactions on Smart Grid 11 (2), p. 1343ā1355. Cited by: §I-B. [12] Z. Zhu, Z. Hu, K. W. Chan, S. Bu, B. Zhou, and S. Xia (2023) Reinforcement learning in deregulated energy market: a comprehensive review. Applied Energy 329, p. 120212. External Links: ISSN 0306-2619, Document, Link Cited by: §I-B. [13] F. Hu, Y. Zhao, Y. Yu, C. Zhang, Y. Lian, C. Huang, and Y. Li (2025) Strategic bidding with price-quantity pairs based on deep reinforcement learning considering competitorsā behaviors. Applied Energy 391, p. 125874. External Links: ISSN 0306-2619, Document, Link Cited by: §I-B. [14] H. Zhang, G. Tsaousoglou, S. Zhan, K. Kok, and N. G. Paterakis (2025) Taming deep reinforcement learning agents with pricing mechanism: validation in power distribution systems. Energy and AI, p. 100635. Cited by: §I-B. [15] N. Harder, R. Qussous, and A. Weidlich (2023) Fit for purpose: modeling wholesale electricity markets realistically with multi-agent deep reinforcement learning. Energy and AI 14, p. 100295. External Links: ISSN 2666-5468, Document, Link Cited by: §I-B. [16] Y. Du, F. Li, H. Zandi, and Y. Xue (2021) Approximating nash equilibrium in day-ahead electricity market bidding with multi-agent deep reinforcement learning. Journal of Modern Power Systems and Clean Energy 9 (3), p. 534ā544. External Links: ISSN 2196-5625, Document, Link Cited by: §I-B. [17] Z. Pan and Z. Jing (2025) Decision-making and cost models of generation company agents for supporting future electricity market mechanism design based on agent-based simulation. Applied Energy 391, p. 125881. External Links: ISSN 0306-2619, Document, Link Cited by: §I-B. [18] T. Klein (2021) Autonomous algorithmic collusion: qālearning under sequential pricing. The RAND Journal of Economics 52 (3), p. 538ā558. External Links: ISSN 1756-2171, Link, Document Cited by: §I-B. [19] T. Werner (2021) Algorithmic and human collusion. SSRN Electronic Journal. External Links: ISSN 1556-5068, Link, Document Cited by: §I-B. [20] D. Esmaeili Aliabadi and K. Chan (2022) The emerging threat of artificial intelligence on competition in liberalized electricity markets: a deep q-network approach. Applied Energy 325, p. 119813. External Links: ISSN 0306-2619, Document, Link Cited by: §I-B, §I-C, §IV-A. [21] M. S. Ćvila, R. Ebrahimy, and G. Tsaousoglou (2025) How inefficient can an electricity market be?. In 2025 IEEE Kiel PowerTech, p. 1ā6. Cited by: §I-C. [22] J. Miklós-Thal (2011) Optimal collusion under cost asymmetry. Economic Theory 46 (1), p. 99ā125. External Links: ISSN 09382259, 14320479, Link Cited by: §IV-A.