SearcharxivSearch

arXiv subjects

Martin Bichler

Publications and source records attributed to Martin Bichler.

At least 19 recordsLinked to original sources

Convergence of Stochastic First-Order Algorithms in Bertrand Competition Under Incomplete Information

Autonomous pricing agents are widely deployed in online marketplaces, making algorithmic pricing a prominent application of multi-agent learning. Experimental studies often report collusive outcomes, but these findings typically rely on Q-learning in complete-information environments and lack rigorous convergence guarantees. In this paper, we study the stochastic learning dynamics of Regularized Robbins-Monro (RRM) algorithms in a Bayesian Bertrand competition with private costs. We show that this setting violates standard stability conditions, including monotonicity and the Minty variational inequality, rendering classical convergence results for gradient-based learning inapplicable. Despite this, we prove that Euclidean RRM algorithms converge almost surely to the unique, efficient Bayes-Nash equilibrium within a finite-dimensional approximation of the strategy space. By analyzing symmetric piecewise-linear pricing strategies in a duopoly, we explicitly construct a global Lyapunov function for the projected primal dynamics and establish global asymptotic stability of the equilibrium. Our analysis yields rigorous convergence guarantees for stochastic first-order learning algorithms in Bayesian Bertrand competition and provides a principled counterpoint to widespread claims of algorithmic collusion.

cs.GT

Algorithmic Predation: Equilibrium Analysis in Dynamic Oligopolies with Smooth Market Sharing

Predatory pricing -- where a firm strategically lowers prices to undermine competitors -- is a contentious topic in dynamic oligopoly theory, with scholars debating practical relevance and the existence of predatory equilibria. Although finite-horizon dynamic models have long been proposed to capture the strategic intertemporal incentives of oligopolists, the existence and form of equilibrium strategies in settings that allow for firm exit (drop-outs following loss-making periods) have remained an open question. We focus on the seminal dynamic oligopoly model by Selten (1965) that introduces the subgame perfect equilibrium and analyzes smooth market sharing. Equilibrium can be derived analytically in models that do not allow for dropouts, but not in models that can lead to predatory pricing. In this paper, we leverage recent advances in deep reinforcement learning to compute and verify equilibria in finite-horizon dynamic oligopoly games. Our experiments reveal two key findings: first, state-of-the-art deep reinforcement learning algorithms reliably converge to equilibrium in both perfect- and imperfect-information oligopoly models; second, when firms face asymmetric cost structures, the resulting equilibria exhibit predatory pricing behavior. These results demonstrate that predatory pricing can emerge as a rational equilibrium strategy across a broad variety of model settings. By providing equilibrium analysis of finite-horizon dynamic oligopoly models with drop-outs, our study answers a decade-old question and offers new insights for competition authorities and regulators.

cs.GT

Agentic Markets: Game Dynamics and Equilibrium in Markets with Learning Agents

Autonomous and learning agents increasingly participate in markets - setting prices, placing bids, ordering inventory. Such agents are not just aiming to optimize in an uncertain environment; they are making decisions in a game-theoretical environment where the decision of one agent influences the profit of other agents. While game theory usually predicts outcomes of strategic interaction as an equilibrium, it does not capture how repeated interaction of learning agents arrives at a certain outcome. This article surveys developments in modeling agent behavior as dynamical systems, with a focus on projected gradient and no-regret learning algorithms. In general, learning in games can lead to all types of dynamics, including convergence to equilibrium, but also cycles and chaotic behavior. It is important to understand when we can expect efficient equilibrium in automated markets and when this is not the case. Thus, we analyze when and how learning agents converge to an equilibrium of a market game, drawing on tools from variational inequalities and Lyapunov stability theory. Special attention is given to the stability of projected dynamics and the convergence to equilibrium sets as limiting outcomes. Overall, the paper provides mathematical foundations for analyzing stability and convergence in agentic markets driven by autonomous, learning agents.

cs.GT

Distributionally Robust Contract Design with Deferred Inspection

We study a robust contract design problem with deferred inspection, in which a principal allocates a scarce resource to an agent, observes the agent's realized outcome ex post at negligible cost, and conditions transfers on this information through rewards. The principal faces ambiguity about the agent's value distribution and seeks to maximize worst-case expected revenue subject to incentive compatibility and limited liability. In contrast to existing work on inspection mechanisms, which relies on common-prior assumptions, we adopt a distributionally robust approach based on moment information. Our main contribution is a complete characterization of the robust contract design problem with a single agent. When the ambiguity set is defined by the first moment, we identify a robustly optimal contract with a concave allocation rule and a linear payment rule. We further show that robustness does not uniquely pin down transfers: we construct a Pareto robustly optimal contract that preserves the same allocation while extracting maximal feasible payments from all types, yielding strictly higher expected revenue under non-worst-case distributions. We also derive structural results for multi-agent extensions. For ambiguity sets defined by the first $N$ moments, we show that robust optimality requires aggregate payments to be lower bounded by a multi-dimensional polynomial of degree $N$. However, unlike the single-agent case, robust multi-agent mechanisms are substantially more complex: dominant-strategy incentive compatibility becomes necessary, simple monotone mechanisms are no longer tractable, and worst-case distributions may involve correlated types or degenerate to a Dirac distribution at the mean. These results highlight a sharp contrast between robust contract design and robust multi-agent mechanism design with inspection.

econ.TH

Algorithmic Pricing and Algorithmic Collusion

The rise of algorithmic pricing in online retail platforms has attracted significant interest in how autonomous software agents interact under competition. This article explores the potential emergence of algorithmic collusion - supra-competitive pricing outcomes that arise without explicit agreements - as a consequence of repeated interactions between learning agents. Most of the literature focuses on oligopoly pricing environments modeled as repeated Bertrand competitions, where firms use online learning algorithms to adapt prices over time. While experimental research has demonstrated that specific reinforcement learning algorithms can learn to maintain prices above competitive equilibrium levels in simulated environments, theoretical understanding of when and why such outcomes occur remains limited. This work highlights the interdisciplinary nature of this challenge, which connects computer science concepts of online learning with game-theoretical literature on equilibrium learning. We examine implications for the Business & Information Systems Engineering (BISE) community and identify specific research opportunities to address challenges of algorithmic competition in digital marketplaces.

cs.GT

Characterizing the Convergence of Game Dynamics via Potentialness

Understanding the convergence landscape of multi-agent learning is a fundamental problem of great practical relevance in many applications of artificial intelligence and machine learning. While it is known that learning dynamics converge to Nash equilibrium in potential games, the behavior of dynamics in many important classes of games that do not admit a potential is poorly understood. To measure how ''close'' a game is to being potential, we consider a distance function, that we call ''potentialness'', and which relies on a strategic decomposition of games introduced by Candogan et al. (2011). We introduce a numerical framework enabling the computation of this metric, which we use to calculate the degree of ''potentialness'' in generic matrix games, as well as (non-generic) games that are important in economic applications, namely auctions and contests. Understanding learning in the latter games has become increasingly important due to the wide-spread automation of bidding and pricing with no-regret learning algorithms. We empirically show that potentialness decreases and concentrates with an increasing number of agents or actions; in addition, potentialness turns out to be a good predictor for the existence of pure Nash equilibria and the convergence of no-regret learning algorithms in matrix games. In particular, we observe that potentialness is very low for complete-information models of the all-pay auction where no pure Nash equilibrium exists, and much higher for Tullock contests, first-, and second-price auctions, explaining the success of learning in the latter. In the incomplete-information version of the all-pay auction, a pure Bayes-Nash equilibrium exists and it can be learned with gradient-based algorithms. Potentialness nicely characterizes these differences to the complete-information version.

cs.GT

Exploring Competitive and Collusive Behaviors in Algorithmic Pricing with Deep Reinforcement Learning

Nowadays, a significant share of the business-to-consumer sector is based on online platforms like Amazon and Alibaba and uses AI for pricing strategies. This has sparked debate on whether pricing algorithms may tacitly collude to set supra-competitive prices without being explicitly designed to do so. Our study addresses these concerns by examining the risk of collusion when Reinforcement Learning (RL) algorithms are used to decide on pricing strategies in competitive markets. Prior research in this field focused on Tabular Q-learning (TQL) and led to opposing views on whether learning-based algorithms can result in supra-competitive prices. Building on this, our work contributes to this ongoing discussion by providing a more nuanced numerical study that goes beyond TQL, additionally capturing off- and on- policy Deep Reinforcement Learning (DRL) algorithms, two distinct families of DRL algorithms that recently gained attention for algorithmic pricing. We study multiple Bertrand oligopoly variants and show that algorithmic collusion depends on the algorithm used. In our experiments, we observed that TQL tends to exhibit higher collusion and price dispersion. Moreover, it suffers from instability and disparity, as agents with higher learning rates consistently achieve higher profits, and it lacks robustness in state representation, with pricing dynamics varying significantly based on information access. In contrast, DRL algorithms, such as PPO and DQN, generally converge to lower prices closer to the Nash equilibrium. Additionally, we show that when pre-trained TQL agents interact with DRL agents, the latter quickly outperforms the former, highlighting the advantages of DRL in pricing competition. Lastly, we find that competition between heterogeneous DRL algorithms, such as PPO and DQN, tends to reduce the likelihood of supra-competitive pricing.

econ.GN

Semicoarse Correlated Equilibria and LP-Based Guarantees for Gradient Dynamics in Normal-Form Games

Projected gradient ascent is known to satisfy no-external regret as a learning algorithm. However, recent empirical work shows that projected gradient ascent often finds the Nash equilibrium in settings beyond two-player zero-sum interactions or potential games, including those where the set of coarse correlated equilibria is very large. We show that gradient ascent in fact satisfies a stronger class of linear $\Phi$-regret in normal-form games; resulting in a refined solution concept which we dub semicoarse correlated equilibria. Our theoretical analysis of the discretised Bertrand competition mirrors those recently established for mean-based learning in first-price auctions. With at least two firms of lowest marginal cost, Nash equilibria emerge as the only semicoarse equilibria under concavity conditions on firm profits. In first-price auctions, the granularity of the bid space affects semicoarse equilibria, but finer granularity for lower bids also induces convergence to Nash equilibria. Unlike previous work that aims to prove convergence to a Nash equilibrium that often relies on epoch based analysis and probability theoretic machinery, our LP-based duality approach enables a simple and tractable analysis of equilibrium selection under gradient-based learning.

cs.GT

Online Optimization Algorithms in Repeated Price Competition: Equilibrium Learning and Algorithmic Collusion

This paper investigates whether online learning algorithms in pricing produce competitive outcomes or tacit collusion. This issue has drawn considerable attention from competition regulators as algorithmic pricing becomes more common in digital markets. Understanding when such algorithms lead to equilibrium or supra-competitive prices is critical for buyers, sellers, and policymakers. We study the behavior of multi-armed bandit (MAB) online learning algorithms in repeated price competition. These algorithms require little information to learn, making them realistic models of automated pricing. Our analysis shows that mean-based algorithms, a special variant of online learning algorithms, converge to correlated rationalizable actions. In the Bertrand environments considered, this implies convergence to the Nash equilibrium or adjacent prices. Numerical experiments reveal that most MAB algorithms, including those that are not mean-based, also converge. We observe supra-competitive prices only in specific cases where all sellers implement the same symmetric version of certain algorithms, such as UCB. This effect diminishes as the number of competitors increases. Our results suggest that, even in a stylized repeated Bertrand competition, sustained supra-competitive prices may be less of a concern when independent agents use different online learning algorithms. Our insights are relevant for regulators and managers considering the use of algorithmic pricing algorithms.

cs.GT

Algorithmic Collusion in Dynamic Pricing with Deep Reinforcement Learning

Nowadays, a significant share of the Business-to-Consumer sector is based on online platforms like Amazon and Alibaba and uses Artificial Intelligence for pricing strategies. This has sparked debate on whether pricing algorithms may tacitly collude to set supra-competitive prices without being explicitly designed to do so. Our study addresses these concerns by examining the risk of collusion when Reinforcement Learning algorithms are used to decide on pricing strategies in competitive markets. Prior research in this field focused on Tabular Q-learning (TQL) and led to opposing views on whether learning-based algorithms can lead to supra-competitive prices. Our work contributes to this ongoing discussion by providing a more nuanced numerical study that goes beyond TQL by additionally capturing off- and on-policy Deep Reinforcement Learning (DRL) algorithms. We study multiple Bertrand oligopoly variants and show that algorithmic collusion depends on the algorithm used. In our experiments, TQL exhibits higher collusion and price dispersion phenomena compared to DRL algorithms. We show that the severity of collusion depends not only on the algorithm used but also on the characteristics of the market environment. We further find that Proximal Policy Optimization appears to be less sensitive to collusive outcomes compared to other state-of-the-art DRL algorithms.

econ.GN

Challenges in Finding Stable Price Zones in European Electricity Markets: Aiming to Square the Circle?

The European day-ahead electricity market is split into multiple bidding zones with a uniform price. The increase in renewables leads to a growing number of interventions in the generation of energy sources and increasing redispatch costs. To ensure efficient congestion management, the EU Commission mandated a Bidding Zone Review (BZR) to reevaluate the configuration of European bidding zones. An integral part of this process was a locational marginal pricing study. Based on these prices, alternative bidding zone configurations were proposed. These bidding zones shall be stable and robust over time. For Germany, four configurations were suggested. We analyzed the proposed configurations considering different clustering algorithms and periods based on the publicly released data set in the context of the BZR, and found that the configurations do not reduce the price standard deviations within zones much, and the average prices across zones are similar. Other configurations identified based on clustering the prices lead to lower price variance but they are not geographically coherent. Independent of the clustering features and algorithms used, the resulting clusters are not stable over time. Interestingly, the effect of a split on prices would be low based on an analysis of the BZR data set.

econ.GN

Zonal vs. Nodal Pricing: An Analysis of Different Pricing Rules in the German Day-Ahead Market

The European electricity market is based on large pricing zones with a uniform day-ahead price. The energy transition leads to changes in supply and demand and increasing redispatch costs. In an attempt to ensure efficient market clearing and congestion management, the EU Commission has mandated the Bidding Zone Review (BZR) to reevaluate the configuration of European bidding zones. Based on a unique data set published in the context of the BZR for the target year 2025, we analyze the short-run effects of various pricing rules for the German-Luxembourgish bidding zone. We compare market clearing and pricing for different zonal models, including their generation and redispatch costs. In numerical experiments with this dataset, the differences in the average prices in different zones are low. The total costs across different configurations are similar and the reduction of standard deviations in prices is also small. This might be different with other load and generation scenarios, but the BZR data is important as it was created to make a decision about splits of the existing bidding zones. We can replicate several results from the BZR study, except the large cost savings when moving from one to two price zones in Germany and Luxembourg. In addition to the four zonal configurations analyzed in the BZR study, we compare these against a nodal pricing system. While the total cost savings after introducing zonal splits were less than 1%, nodal pricing led to savings of 5-6%. We also evaluate differences of nodal pricing rules with respect to the necessary uplift payments, which is relevant in the context of the discussion on non-uniform pricing in the EU. While the study focuses on Germany, the analysis is relevant beyond and feeds into the broader discussion about pricing rules in non-convex markets.

econ.GN

On the Uniqueness of Bayesian Coarse Correlated Equilibria in Standard First-Price and All-Pay Auctions

We study the Bayesian coarse correlated equilibrium (BCCE) of continuous and discretised first-price and all-pay auctions under the standard symmetric independent private-values model. Our study is motivated by the question of how the canonical Bayes-Nash equilibrium (BNE) of the auction relates to the outcomes learned by buyers utilising no-regret algorithms. Numerical experiments show that in two buyer first-price auctions the Wasserstein-$2$ distance of buyers' marginal bid distributions decline as $O(1/n)$ in the discretisation size in instances where the prior distribution is concave, whereas all-pay auctions exhibit similar behaviour without prior dependence. To explain this convergence to a near-equilibrium, we study uniqueness of the BCCE of the continuous auction. Our uniqueness results translate to provable convergence of deterministic self-play to a near equilibrium outcome in these auctions. In the all-pay auction, we show that independent of the prior distribution there is a unique BCCE with symmetric, differentiable, and increasing bidding strategies, which is equivalent to the unique strict BNE. In the first-price auction, we need stronger conditions. Either the prior is strictly concave or the learning algorithm has to be restricted to strictly increasing strategies. Without such strong assumptions, no-regret algorithms can end up in low-price pooling strategies. This is important because it proves that in repeated first-price auctions such as in display ad actions, algorithmic collusion cannot be ruled out without further assumptions even if all bidders rely on no-regret algorithms.

cs.GT

Equilibrium Computation in Multi-Stage Auctions and Contests

We compute equilibrium strategies in multi-stage games with continuous signal and action spaces as they are widely used in the management sciences and economics. Examples include sequential sales via auctions, multi-stage elimination contests, and Stackelberg competitions. In sequential auctions, analysts performing equilibrium analysis are required to derive not just single bids but bid functions for all possible signals or values that a bidder might have in multiple stages. Due to the continuity of the signal and action spaces, these bid functions come from an infinite dimensional space. While such models are fundamental to game theory and its applications, equilibrium strategies are rarely known. The resulting system of non-linear differential equations is considered intractable for all but elementary models. This has been limiting progress in game theory and is a barrier to its adoption in the field. We show that Deep Reinforcement Learning and self-play can learn equilibrium bidding strategies for various multi-stage games. We find equilibrium in models that have not yet been explored analytically and new asymmetric equilibrium bid functions for established models of sequential auctions. The verification of equilibrium is challenging in such games due to the continuous signal and action spaces. We introduce a verification algorithm and prove that the error of this verifier decreases when considering Lipschitz continuous strategies with increasing levels of discretization and sample sizes.

cs.GT

Solving Large-Scale Electricity Market Pricing Problems in Polynomial Time

Electricity market operators worldwide use mixed-integer linear programming to solve the allocation problem in wholesale electricity markets. Prices are typically determined based on the duals of relaxed versions of this optimization problem. The resulting outcomes are efficient, but market operators must pay out-of-market uplifts to some market participants and incur a considerable budget deficit that was criticized by regulators. As the share of renewables increases, the number of market participants will grow, leading to larger optimization problems and runtime issues. At the same time, non-convexities will continue to matter e.g., due to ramping constraints of the generators required to address the variability of renewables or non-convex curtailment costs. We draw on recent theoretical advances in the approximation of competitive equilibrium to compute allocations and prices in electricity markets using convex optimization. The proposed mechanism promises approximate efficiency, no budget deficit, and computational tractability. We present experimental results for this new mechanism in the context of electricity markets, and compare the runtimes, the average efficiency loss of the method, and the uplifts paid with standard pricing rules. We find that the computations with the new algorithm are considerably fast for relevant problem sizes. In general, the computational advantages come at the cost of efficiency losses and a price markup for the demand side. Interestingly, both are small with realistic problem instances. Importantly, the market operator does not incur a budget deficit and the uplifts paid to market participants are significantly lower compared to standard pricing rules.

cs.GT

Revenue in First- and Second-Price Display Advertising Auctions: Understanding Markets with Learning Agents

The transition of display ad exchanges from second-price auctions (SPA) to first-price auctions (FPA) has raised questions about its impact on revenue. Auction theory predicts the revenue equivalence between these two auction formats. However, display ad auctions are different from standard models in auction theory. First, automated bidding agents cannot easily derive equilibrium strategies in FPA because information regarding competitors is not readily available. Second, due to principal-agent problems, bidding agents typically maximize return-on-investment (ROI), not payoff. The literature on learning agents for real-time bidding is growing because of the practical relevance of this area; most research has found that learning agents do not converge to an equilibrium. Specifically, research on algorithmic collusion in display ad auctions has argued that FPA can induce symmetric Q-learning agents to tacitly collude, resulting in bids below equilibrium, leading to lower revenue compared to the SPA. Whether bids are in equilibrium cannot easily be determined from field data since the underlying values of bidders are unknown. In this paper, we draw on analytical modeling and numerical experiments and explore the convergence behavior of widespread online learning algorithms in both complete and incomplete information models. Contrary to prior results, we show that there are no systematic deviations from equilibrium behavior. We also explore the differences in revenue of the FPA and SPA, which have not been done for utility functions relevant to this domain, such as ROI. We show that learning algorithms also converge to equilibrium. Still, revenue equivalence does not hold, indicating that collusion may not be the explanation for lower revenue with FPA, and the change in auction format might have had substantial and non-obvious consequences for ad exchanges and advertisers.

cs.GT

On the Convergence of Learning Algorithms in Bayesian Auction Games

Equilibrium problems in Bayesian auction games can be described as systems of differential equations. Depending on the model assumptions, these equations might be such that we do not have a rigorous mathematical solution theory. The lack of analytical or numerical techniques with guaranteed convergence for the equilibrium problem has plagued the field and limited equilibrium analysis to rather simple auction models such as single-object auctions. Recent advances in equilibrium learning led to algorithms that find equilibrium under a wide variety of model assumptions. We analyze first- and second-price auctions where simple learning algorithms converge to an equilibrium. The equilibrium problem in auctions is equivalent to solving an infinite-dimensional variational inequality (VI). Monotonicity and the Minty condition are the central sufficient conditions for learning algorithms to converge to an equilibrium in such VIs. We show that neither monotonicity nor pseudo- or quasi-monotonicity holds for the respective VIs. The second-price auction's equilibrium is a Minty-type solution, but the first-price auction is not. However, the Bayes--Nash equilibrium is the unique solution to the VI within the class of uniformly increasing bid functions, which ensures that gradient-based algorithms attain the equilibrium in case of convergence, as also observed in numerical experiments.

cs.GT

Pricing Optimal Outcomes in Coupled and Non-Convex Markets: Theory and Applications to Electricity Markets

According to the fundamental theorems of welfare economics, any competitive equilibrium is Pareto efficient. Unfortunately, competitive equilibrium prices only exist under strong assumptions such as perfectly divisible goods and convex preferences. In many real-world markets, participants have non-convex preferences and the allocation problem needs to consider complex constraints. Electricity markets are a prime example, but similar problems appear in many real-world markets, which has led to a growing literature in market design. Power markets use heuristic pricing rules based on the dual of a relaxed allocation problem today. With increasing levels of renewables, these rules have come under scrutiny as they lead to high out-of-market side-payments to some participants and to inadequate congestion signals. We show that existing pricing heuristics optimize specific design goals that can be conflicting. The trade-offs can be substantial, and we establish that the design of pricing rules is fundamentally a multi-objective optimization problem addressing different incentives. In addition to traditional multi-objective optimization techniques using weighing of individual objectives, we introduce a novel parameter-free pricing rule that minimizes incentives for market participants to deviate locally. Our theoretical and experimental findings show how the new pricing rule capitalizes on the upsides of existing pricing rules under scrutiny today. It leads to prices that incur low make-whole payments while providing adequate congestion signals and low lost opportunity costs. Our suggested pricing rule does not require weighing of objectives, it is computationally scalable, and balances trade-offs in a principled manner, addressing an important policy issue in electricity markets.

cs.GT