Searcharxiv⌕ Search

arXiv subjects

Janusz M. Meylahn

Publications and source records attributed to Janusz M. Meylahn.

7 recordsLinked to original sources

Auditing Algorithmic Collusion from Strategy Graphs

Detecting algorithmic collusion is challenging because regulators often have limited access to firms' algorithms, training data, and market information. We study an intermediate-information regime in which an auditor can query firms' frozen pricing policies and construct the induced strategy graph. Using a complete characterization of Nash equilibria in a repeated pricing game, we identify graph-theoretic features of strategy graphs that are associated with collusive reward-and-punishment schemes, including maximum betweenness, attractor in-degree, and average path length. We then test these metrics on policies learned by decentralized Q-learning and the Q-learning algorithm of Calvano et al. (2020). We find that especially the maximum betweenness and attractor in-degree are strongly correlated with the standard profit-based Collusion Index. Importantly, the proposed metrics rely only on the unlabeled topology of strategy graphs and require neither price histories, demand estimates, nor competitive and monopoly benchmarks. Our results suggest that the structure of frozen pricing policies contains robust signals of collusion among reinforcement learning algorithms and provides a promising basis for auditing algorithmic pricing systems under limited information.

econ.TH↗

Equilibrium stability as a driver of cooperation among Q-learners

Algorithmic collusion among pricing algorithms has raised concerns about sustained supra-competitive prices and their implications for social welfare. Existing work has largely focused on the probability that reinforcement-learning algorithms converge to cooperative strategies, typically under the assumption that exploration vanishes over time. Motivated by the observation that algorithms deployed in practice are likely to continue exploring in order to remain adaptive to changing environments, we study learning dynamics under constant exploration. In this setting, the relevant question is no longer whether an algorithm converges to a particular strategy profile, but rather what fraction of time the algorithms spend playing cooperative strategies. Even in the benchmark case of the repeated Prisoner's Dilemma with one-period memory, this yields high-dimensional stochastic learning dynamics, for which a complete analytic treatment is intractable. We show that cooperative strategies can be dominant in this time-averaged sense and derive a boundary predicting when such dominance arises, based on the expected dynamics of the Q-learning process. Extensive simulations show that this boundary is a strong predictor for non-defection-dominated behaviour under epsilon-greedy Q-learning.

cs.MA↗

Beyond the Independence Assumption: Finite-Sample Guarantees for Deep Q-Learning under $τ$-Mixing

Finite-sample analyses of deep Q-learning typically treat replayed data as independent, even though it is sampled from temporally dependent state-action trajectories. We study the Deep Q-networks (DQN) algorithm under explicit dependence by modelling the minibatches used for updating the network as $τ$-mixing. We show that this assumption holds under certain dependence conditions on the underlying trajectories and the mechanism used to sample minibatches. Building on this observation, we extend statistical analyses of DQN with fully connected ReLU architectures to dependent data. We formulate each update as a nonparametric regression problem with $τ$-mixing observations and derive finite-sample risk bounds under this dependence structure. Our results show that temporal dependence leads to a degradation in the statistical rate by inducing an additional dimensionality penalty in the rate exponent, reflecting the reduced effective sample size of $τ$-mixing data. Moreover, we derive the sample complexity of DQN under $tau$-mixing from these risk bounds. Finally, we empirically demonstrate on standard Gymnasium environments that the independence assumption is systematically violated and that replay sampling yields approximately exponentially decaying correlations, supporting our theoretical framework.

stat.ML↗

How social reinforcement learning can lead to metastable polarisation and the voter model

Previous explanations for the persistence of polarization of opinions have typically included modelling assumptions that predispose the possibility of polarization (i.e., assumptions allowing a pair of agents to drift apart in their opinion such as repulsive interactions or bounded confidence). An exception is a recent simulation study showing that polarization is persistent when agents form their opinions using social reinforcement learning. Our goal is to highlight the usefulness of reinforcement learning in the context of modeling opinion dynamics, but that caution is required when selecting the tools used to study such a model. We show that the polarization observed in the model of the simulation study cannot persist indefinitely, and exhibits consensus asymptotically with probability one. By constructing a link between the reinforcement learning model and the voter model, we argue that the observed polarization is metastable. Finally, we show that a slight modification in the learning process of the agents changes the model from being non-ergodic to being ergodic. Our results show that reinforcement learning may be a powerful method for modelling polarization in opinion dynamics, but that the tools (objects to study such as the stationary distribution, or time to absorption for example) appropriate for analysing such models crucially depend on their properties (such as ergodicity, or transience). These properties are determined by the details of the learning process and may be difficult to identify based solely on simulations.

physics.soc-ph↗

Two-community noisy Kuramoto model with general interaction strengths: Part I

We generalize the study of the noisy Kuramoto model, considered on a network of two interacting communities, to the case where the interaction strengths within and across communities are taken to be different in general. By developing a geometric interpretation of the self-consistency equations, we are able to separate the parameter space into ten regions in which we identify the maximum number of solutions in the steady state. Furthermore, we prove that in the steady-state only the angles 0 and $π$ are possible between the average phases of the two communities and derive the solution boundary for the unsynchronized solution. Lastly, we identify the equivalence class relation in the parameter space corresponding to the symmetrically synchronized solution.

math-ph↗

Properties of additive functionals of Brownian motion with resetting

We study the distribution of additive functionals of reset Brownian motion, a variation of normal Brownian motion in which the path is interrupted at a given rate and placed back to a given reset position. Our goal is two-fold: (1) For general functionals, we derive a large deviation principle in the presence of resetting and identify the large deviation rate function in terms of a variational formula involving large deviation rate functions without resetting. (2) For three examples of functionals (positive occupation time, area and absolute area), we investigate the effect of resetting by computing distributions and moments, using a formula that links the generating function with resetting to the generating function without resetting.

math.PR↗

Large deviations for Markov processes with resetting

Markov processes restarted or reset at random times to a fixed state or region in space have been actively studied recently in connection with random searches, foraging, and population dynamics. Here we study the large deviations of time-additive functions or observables of Markov processes with resetting. By deriving a renewal formula linking generating functions with and without resetting we are able to obtain the rate function of such observables, characterizing the likelihood of their fluctuations in the long-time limit. We consider as an illustration the large deviations of the area of the Ornstein-Uhlenbeck process with resetting. Other applications involving diffusions, random walks, and jump processes with resetting or catastrophes are discussed.

cond-mat.stat-mech↗