Searcharxiv⌕ Search

arXiv · 2610.00619

Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games

Abstract

In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that deters deviations and provides behavioural evidence of collusion. We investigate this mechanism in a two-player, finite-horizon Almgren-Chriss liquidation game. Independent proximal policy optimisation agents with access to within-episode price and action histories achieve costs below the Nash benchmark. We identify a profitable deviation by training against the mean learned liquidation schedule, then impose its first trade on one of the original agents. The opponent responds by accelerating liquidation. This response more than offsets the deviator's gain in every run and both player roles, while leaving the punisher's average payoff materially unchanged relative to not punishing under the same deviation. The punisher imposes greater losses on the deviator while preserving its own average payoff, despite the availability of more profitable, less punitive liquidation plans. Matching deviations and subsequent additional selling rise and later decline during training, while final policies retain an effective punitive response. We formalise two checks: whether punishment outweighs the gain from deviating, and whether the change in trading behaviour is large enough to account for the loss imposed. Both checks hold for the tested deviation. Together, these findings provide behavioural and economic evidence supporting a collusive interpretation of the learned supra-competitive outcomes.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Christos Spyridon Koulouris, Carlo Campajola. 2026-09-30. Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games. https://arxiv.org/abs/2610.00619

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Social welfare and price discovery in double auction markets

The tendency of the double auction mechanism to drive prices to competitive equilibrium has been well documented in laboratory experiments, but the phenomenon has lacked a theoretical explanation. This paper studies dynamic double auctions in a pure exchange economy where agents bid their indifference prices implied by their current holdings and preferences. We show that Walras equilibria coincide with the fixed points of the double auction and that repeated double auctions generate bounded sequences of allocations and prices whose cluster points are Walras equilibria with transfers.

q-fin.TR↗

Learning Optimal Liquidation with Closing Auctions

We study liquidation when continuous trading is followed by a closing auction. The trader first sells through a limit-order book, then submits signed auction schedules to adjust the remaining inventory. A projected clearing price signal and intermediate auction feedback connect the two phases. We compare deep Q-network (DQN) policies with projected deep deterministic policy gradient (DDPG), twin delayed deep deterministic policy gradient (TD3) and soft actor-critic (SAC) policies, using synthetic rough Heston prices and historical midprice paths within a simulated market. Policies are selected and evaluated by inventory-penalized implementation shortfall, separately from a weighted training objective. They achieve lower inventory-penalized shortfall than the stylized Avellaneda-Stoikov (AS) and time-weighted average price (TWAP) references, and matched synthetic comparisons show that auction access is useful for all four learners. We furthermore find that dense auction credit improves learning; the clearing forecast is informative, but its incremental decision value is learner-dependent.

q-fin.TR↗

A Generalized Langevin Model of Latent Liquidity and Concave Price Impact

We model market impact as the response to submitted order flow net of counterflow from latent traders, activated when price displacements from the level that would prevail without the order exceed individual thresholds. Order flow depletes this pool, and a generalized Langevin equation governs its recovery over several time scales. Its memory kernels are finite sums of exponentials, so its Markovian lift is exact rather than an approximation. For an undepleted pool, aggregation under explicit assumptions on individual trading responses yields an intermediate square-root regime between linear small- and large-order limits, without imposing a square-root impact law. Scaling thresholds and responses with price noise makes impact in this regime proportional to volatility, and thresholds that grow with the execution horizon make it independent of duration. With constant displayed depth, expected round-trip costs are nonnegative under the log-price convention, independently of the memory. Numerical experiments show that depletion narrows the square-root range and that memory spectra producing similar single-order impacts can respond differently after substantial prior trading. Calibration to market data is left to a companion paper.

q-fin.TR↗