Source-linked AI summary

AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion

Jakub Seredyński, Georgios Tsaousoglou

arXiv:2608.26896v1cs.AIcs.GTcs.MAeess.SY

TL;DR

The paper investigates whether autonomous learning-based bidding agents can generate tacit collusion in electricity markets. It models the market as a repeated game with imperfect public monitoring, uses multi-agent reinforcement learning, and evaluates emergent behavior with multiple criteria. Experiments find realistic cases where agents sustain supra-competitive outcomes supportive of tacit-collusion indicators without being instructed to collude.

  • Problem

    Existing collusion tests based mainly on Nash-equilibrium profit comparisons are difficult to apply to realistic electricity markets, while supracompetitive profits alone are insufficient evidence.

  • Method

    The paper models strategic bidding with multi-agent reinforcement learning and evaluates emergent joint strategies using multiple tacit-collusion criteria.

  • Results

    Under some conditions, particularly binding network constraints, agents learn to sustain supra-competitive outcomes satisfying several independent tacit-collusion indicators without collusive instructions.

  • Takeaways & Limitations

    The results indicate that tacit collusion is a realistic danger in algorithmic electricity markets and warrants further study of learning mechanisms and market-design countermeasures.

  • Takeaways & Limitations

    The authors caution that their results should not be taken as conclusive.

Abstract

from arXiv · show

As electricity market participants increasingly adopt learning-based agents for their bidding strategies, electricity markets are becoming algorithmic. Evidence from algorithmic markets in other domains shows that tacit collusion can arise purely through independent learning. Moreover, electricity markets are typically oligopolistic and feature repeated interaction among a small number of participants, making them structurally susceptible to non-competitive behavior. In the face of these observations, this paper investigates the hypothesis that tacit collusion may emerge in electricity markets where participants' actions are controlled by autonomous learning-based algorithms. We model strategic bidding as a repeated game with imperfect public monitoring, and model the participants' emergent behavior using multi-agent reinforcement learning. We propose a multi-dimensional set of criteria (going beyond profit comparisons against Nash equilibria) to assess whether the resulting behavior constitutes tacit collusion. Our experimental results showcase that such a danger is realistic for electricity markets: there are cases where agents do learn to sustain supra-competitive outcomes that are supportive of tacit collusion indicators, even though the agents were never instructed to collude.

I. INTRODUCTION

The paper asks whether autonomous learning agents can generate tacit collusion in electricity markets, where oligopoly and repeated interaction create structural susceptibility. It argues that existing identification methods are insufficient and motivates a multi-criteria investigation.

  • A. Motivation and Research Question: Tacit collusion can emerge autonomously when independent learning agents coordinate without explicit agreement, potentially raising prices and harming consumers.Prior algorithmic-pricing research found reinforcement-learning agents sustaining supracompetitive prices through trial and error without communication.
  • A. Motivation and Research Question: Electricity markets are especially relevant because they typically combine oligopolistic participation with repeated interaction among a small number of firms.
  • C. Research Gap and Contributions: Existing studies identify collusion partly through profit comparisons with Nash equilibria, an approach that becomes difficult in realistic electricity-market models.
  • C. Research Gap and Contributions: Supracompetitive profits alone neither establish collusive coordination nor rule it out when learned behavior fails to converge to a stage-game Nash equilibrium.
  • C. Research Gap and Contributions: The paper adapts multiple tacit-collusion criteria, models bidding with MARL, and tests how grid constraints, demand, and marginal-cost variance affect emergent collusion.

II. SYSTEM MODEL

The system model represents a single-timeslot, networked electricity market cleared by DC optimal power flow. Generators submit piecewise-linear bids, and dispatch, prices, and profits follow from the market-clearing solution.

  • A. Electricity Market: The market consists of buses with inelastic demand and energy-providing resources, including generators represented across the network.
  • A. Electricity Market: DC-OPF determines dispatch subject to generator, power-balance, phase-angle, power-flow, and transmission-line capacity constraints.
  • A. Electricity Market: Each generator partitions output into segments with bounded quantities and declares per-unit energy costs, producing a piecewise-linear cost function.
  • A. Electricity Market: Market clearing is repeated across timeslots, allowing generators to submit different segment bids and producing timeslot-specific outcomes.

B. Agent-based Market Participation Strategies

Agents represent generators and repeatedly choose bids from private histories, then learn from their own market outcomes under imperfect monitoring. MARL produces an emergent joint bidding strategy profile.

  • B. Agent-based Market Participation Strategies: Agents observe their own dispatch and profit together with locational marginal prices at every network node, but not other agents’ actions or profits.
  • B. Agent-based Market Participation Strategies: The setting is a repeated game with imperfect monitoring because locational marginal prices provide public signals that do not reveal other agents’ actions or profits.
  • B. Agent-based Market Participation Strategies: Each agent seeks a strategy that maximizes cumulative profits over time, with actions determined by its strategy and current history.
  • B. Agent-based Market Participation Strategies: Each agent maps its private history to a bid, while the joint strategy profile captures the bidding behavior that emerges from simultaneous learning.
  • B. Agent-based Market Participation Strategies: MARL is used because each agent faces a non-stationary environment while other agents’ strategies evolve simultaneously and remain unknown.

III. TACIT COLLUSION

The paper evaluates whether an emergent joint strategy is tacitly collusive using three complementary criteria rather than a single behavioral signal.

  • III. TACIT COLLUSION: The three indicators are Punishment of Deviation, Unsustainability under Shortsight, and Short-term Profitability of Deviations.
  • III. TACIT COLLUSION: Together, the criteria assess whether learned joint behavior involves punishment, dependence on forward-looking incentives, and profitable short-run deviations.

A. Punishment of Deviation Criterion

The Punishment of Deviation Criterion tests whether a competitive deviation from the learned joint strategy triggers a temporary punitive response before competitors return to their prior strategies.

  • Punishment of Deviation Criterion: The criterion switches one agent from the learned MARL strategy to an explicitly competitive, short-term profit-maximizing strategy while others retain their observed policies.The deviator’s optimization assumes that other agents continue their previous strategies.
  • Punishment of Deviation Criterion: The deviation problem is formulated as a single-level mixed-integer linear program by replacing the market-clearing problem with KKT conditions and auxiliary binary variables.
  • Punishment of Deviation Criterion: A punitive response consists of competitors lowering mark-ups, thereby reducing the deviator’s profit.
  • Punishment of Deviation Criterion: Immediately after deviation, competitors should submit lower bids, reducing prices and the deviator’s market share.
  • Punishment of Deviation Criterion: After punishment, competitors should gradually return to their pre-deviation strategies while prices recover.

B. Unsustainability under Shortsight Criterion

The Unsustainability under Shortsight Criterion removes two factors that facilitate tacit collusion—emphasis on future payoffs and memory of history—to test whether coordination becomes harder to learn.

  • Unsustainability under Shortsight Criterion: A high discount factor facilitates tacit collusion by increasing the importance of future payoffs relative to immediate ones.
  • Unsustainability under Shortsight Criterion: Memory of past actions and rewards facilitates conditioning behaviour on history.
  • Unsustainability under Shortsight Criterion: Relative to baseline, removing these elements is expected to produce more competitive strategies with lower prices and profits.
  • Unsustainability under Shortsight Criterion: The shortsight test uses a very low discount factor and severely truncated memory of past actions and rewards.
  • Unsustainability under Shortsight Criterion: The criterion concerns collusive strategies that contain profitable unilateral deviations despite agents refraining from them because short-run opportunism may lead to less profitable future outcomes.
  • Unsustainability under Shortsight Criterion: The test uses iterative best response, or diagonalization, by solving the deviation problem successively for each agent to simulate competitive adaptation.
  • Unsustainability under Shortsight Criterion: A positive test requires deviating agents’ profits to increase initially and eventually fall below their profits under the learned strategy as IBR approaches a Nash equilibrium.

IV. EXPERIMENTAL SETTING

The experiments use a controlled seven-bus electricity-market benchmark and screen a full factorial of network, demand, and cost-heterogeneity conditions before behavioural validation.

  • Experimental Setting: The benchmark is a stylised seven-bus market with an oligopolistic supply side, inelastic nodal demand, transmission constraints, and locational marginal prices.
  • Experimental Setting: Demand, marginal costs, and network parameters remain fixed within each scenario, so observed temporal dynamics arise from repeated bidding and market-clearing feedback.
  • Experimental Setting: Grid scenarios enforce the test-system topology and line limits, whereas NoGrid scenarios retain demand and generator data but make transmission constraints non-binding.
  • Experimental Setting: The full factorial design varies Grid versus NoGrid networks, Low/Medium/High demand, and Low/Medium/High marginal-cost heterogeneity, producing 18 environments.The design is 2 × 3 × 3 = 18 market environments.
  • Scenario Selection Methodology: Screening begins from the 18 scenarios and computes indicators over each training run’s final evaluation window after exploratory noise has decayed.
  • Scenario Selection Methodology: The screening combines price and profit markups with price variance, action correlations, action-change correlation, and Best Response Deviation Gain.
  • Scenario Selection Methodology: A positive BRDG indicates that the learned bidding profile contains an unrealised short-run unilateral deviation incentive.
  • Scenario Selection Methodology: Selected cases must be stable, non-boundary outcomes and contain a non-trivial unilateral deviation incentive suitable for behavioural testing.

C. MARL implementation

The MARL implementation compresses each generator’s segmented bid vector into one learned supply-function slope, trains decentralized actors with centralized critics, and evaluates the resulting policies after exploration decays.

  • MARL implementation: Each agent selects a bid vector with one declared marginal cost per segment, giving its action space dimension |S_i|.
  • MARL implementation: A linear supply function maps output levels to declared marginal costs and provides an auxiliary bidding curve distinct from the generator’s true cost function.
  • MARL implementation: The supply-function intercept equals true marginal cost, while the learned nonnegative slope β_i determines the full bid vector through a deterministic mapping.
  • MARL implementation: The parameterization reduces the action space from |S_i| dimensions to one, which may exclude some collusive policies but makes discovered collusion sufficient to validate the broader hypothesis.
  • MARL implementation: Agents are trained with multi-agent TD3 using centralized training and decentralized execution, with critics seeing joint information but actors submitting bids without communication.
  • MARL implementation: Each scenario is trained for 100,000 market instances across 50 episodes, including a 4,000-instance random-action warm-up.
  • MARL implementation: Exploration noise is gradually removed, and screening indicators are computed over the final 10,000 market instances.

V. RESULTS

Across 18 market environments, Grid cases show higher and more structured price markups than NoGrid cases. Three non-boundary Grid cases are selected for behavioural validation because they combine supra-competitive outcomes with varied market conditions and learned dynamics.

  • Grid cases produce substantially higher supra-competitive markups than NoGrid cases, with higher demand and lower marginal-cost heterogeneity associated with higher markups.NoGrid markups are generally modest and lack a clear ordering across demand or cost heterogeneity.
  • Transmission constraints amplify market power and increase sensitivity of market outcomes to demand and cost structure.Elevated markups alone cannot distinguish coordinated behaviour from market power directly arising from network constraints.
  • Three non-boundary Grid cases advance to behavioural validation: Grid-LowDemand-HighCostDiff, Grid-LowDemand-MediumCostDiff, and Grid-MediumDemand-HighCostDiff.The cases differ in demand conditions, cost asymmetry, and learned strategic dynamics while all exhibiting supra-competitive outcomes.

B. Punishment of Deviations

Forced deviations reveal punish-and-forgive responses in both LowDemand Grid scenarios, but not in the MediumDemand-HighCostDiff scenario. Shortsight ablations remove the LowDemand punishment responses while leaving similar on-path bidding levels, separating enforcement from static network-induced market power.

  • Punishment of Deviations: In Grid-LowDemand-HighCostDiff, G1’s forced deviation triggers competitors to lower bidding slopes before returning toward their pre-deviation strategies.This produces a clear punish-and-forgive pattern.
  • Punishment of Deviations: In Grid-LowDemand-MediumCostDiff, G5 and G6 react strongly while G2 remains near its pre-deviation strategy.The combined response is sufficient to penalize the deviator, making G2 non-pivotal to punishment in this case.
  • Punishment of Deviations: Grid-MediumDemand-HighCostDiff shows no comparable coordinated reduction in competitors’ bidding slopes after G5’s deviation.Its reward trajectory likewise provides no evidence of a comparable enforcement response.
  • Shortsight Ablation: Similar on-path bidding levels persist after shortsight ablation, while the off-path enforcement mechanism is removed.This supports a repeated-game component in the two LowDemand scenarios alongside static market power from transmission constraints, local scarcity, and generator pivotality.
  • Shortsight Ablation: Under the shortsight ablation, the baseline punishment-and-forgiveness responses disappear in both LowDemand scenarios.Competitors’ bidding slopes remain largely unchanged after forced deviations, indicating that removing history eliminates the informational basis for conditioning current actions.

D. Iterative Best Responses

Iterative Best Response trajectories test whether learned supra-competitive outcomes remain attractive after unilateral deviations and subsequent competitive adaptation. The two LowDemand scenarios support this criterion, whereas Grid-MediumDemand-HighCostDiff does not and is more plausibly explained by structural market power.

  • Iterative Best Responses: The two LowDemand scenarios combine profitable unilateral deviations with subsequent reward levels below their MARL benchmarks under IBR.This pattern indicates that short-run deviation gains are followed by a less profitable regime under continued best responses.
  • Iterative Best Responses: Grid-MediumDemand-HighCostDiff lacks the combination of profitable deviations and systematically less profitable subsequent IBR rewards.Its IBR trajectories therefore do not support the Short-term Profitability of Deviations criterion.
  • Synthesis of Behavioural Evidence: The two LowDemand learned regimes provide mutually consistent evidence across three behavioural tests for a repeated-game component consistent with Tacit Collusion.Their interpretation is strengthened by the IBR evidence alongside the other behavioural criteria.
  • Synthesis of Behavioural Evidence: Grid-MediumDemand-HighCostDiff remains supra-competitive but shows neither punishment, ablation-sensitive enforcement, nor collusion-consistent IBR dynamics.Its elevated markup is more plausibly attributed to structural market power than Tacit Collusion.
Loading 2608.26896v1…