SearcharxivSearch

arXiv subjects

Andrea Agazzi

Publications and source records attributed to Andrea Agazzi.

At least 19 recordsLinked to original sources

Quantitative Diffusive Limits for Singular Nonlocal Transport

We study the nonlocal continuity equation \[ \partial_t\mu_b =\operatorname{div}\!\left( \mu_b\nabla\log\bigl((I-b^2\Delta)^{-1}\mu_b\bigr) \right) \] on a closed connected Riemannian manifold. For smooth strictly positive initial data, we prove that as $b \to 0$, its global solution converges to heat flow $\mu(t)$ at the sharp, uniform-in-time rate \[ \sup_{t\ge0}\|\mu_b(t)-\mu(t)\|_{L^1}\le Cb^2. \] The key estimate is the uniform dissipation of a $b$-weighted higher-order resolvent energy, which yields exponential relaxation despite the absence of a Wasserstein gradient-flow structure. On the circle, we also analyze the corresponding deterministic $N$-particle dynamics. A weak--strong modulated energy argument gives \[ \mathbb E\!\left[ \sup_{t\ge0}W_1(\mu_b^N(t),\mu_b(t)) \right] \le C(Nb)^{-1/2} \] for iid initialization. Consequently, the choice $b\asymp N^{-1/5}$ approximates heat flow uniformly in time at rate $N^{-2/5}$.

math.AP

Morrey's problem in $\mathbb{R}^{2 \times 4}$ and $\mathbb{R}^{3 \times 3}_\mathrm{sym}$

We find an explicit rank-one convex non-quasiconvex integrand in $\mathbb{R}^{2\times 4}$: to falsify the quasiconvexity inequality, we exhibit a map $\mathbb{T}^4\to \mathbb{R}^2$ with $12$ non-zero Fourier modes. In fact, this map is obtained from a scalar potential, so we also find a rank one convex integrand in $\mathbb{R}^{4\times 4}_\text{sym}$ which is not quasiconvex. These examples are obtained by transpositions and restrictions of Grabovsky's example of a rank-one convex, non-quasiconvex integrand in $\mathbb{R}^{8 \times 2}.$ We also modify \v{S}ver\'{a}k's example to construct a rank-one convex non-quasiconvex integrand in $\mathbb{R}^{3\times 3}_\text{sym}$.

math.AP

Quantitative Gaussian-Process limits of Tensor Programs

We study the infinite-width Gaussian-process limit of random neural networks through the lens of tensor programs, and we provide a quantitative convergence theory in Wasserstein distance. Our main result gives explicit finite-width error bounds, of order inverse square-root of the widths between finite-network executions and their Gaussian-process limits. The framework is architecture-agnostic and covers feed-forward models together with weight-sharing schemes relevant for recurrent and transformer-type architectures.

cs.LG

Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models

We prove pathwise convergence of the layerwise evolution of tokens in a finite-depth, finite-width transformer model with MultiLayer Perceptron (MLP) blocks to a continuous-time stochastic interacting particle system. We also identify the stochastic partial differential equation describing the evolution of the tokens' distribution in this limit and prove propagation of chaos when the number of such tokens is large. The bounds we establish are quantitative and the limits we consider commute. We further prove that the limiting stochastic model displays synchronization by noise and establish exponential dissipation of the interaction energy on average, provided that the common noise is sufficiently coercive relative to the deterministic self-attention drift. We finally characterize the activation functions satisfying the former condition.

math.PR

Token Economy for Fair and Efficient Dynamic Resource Allocation in Congestion Games

Self-interested behavior in sharing economies often leads to inefficient aggregate outcomes compared to a centrally coordinated allocation, ultimately harming users. Yet, centralized coordination removes individual decision power. This issue can be addressed by designing rules that align individual preferences with system-level objectives. Unfortunately, rules based on conventional monetary mechanisms introduce unfairness by discriminating among users based on their wealth. To solve this problem, in this paper, we propose a token-based mechanism for congestion games that achieves efficient and fair dynamic resource allocation. Specifically, we model the token economy as a continuous-time dynamic game with finitely many boundedly rational agents, explicitly capturing their evolutionary policy-revision dynamics. We derive a mean-field approximation of the finite-population game and establish strong approximation guarantees between the mean-field and the finite-population games. This approximation enables the design of integer tolls in closed form that provably steer the aggregate dynamics toward an optimal efficient and fair allocation from any initial condition.

cs.GT

Evolutionary Analysis of Continuous-time Finite-state Mean Field Games with Discounted Payoffs

We consider a class of continuous-time dynamic games involving a large number of players. Each player selects actions from a finite set and evolves through a finite set of states. State transitions occur stochastically and depend on the player's chosen action. A player's single-stage reward depends on their state, action, and the population-wide distribution of states and actions, capturing aggregate effects such as congestion in traffic networks. Each player seeks to maximize a discounted infinite-horizon reward. Existing evolutionary game-theoretic approaches introduce a model for the way individual players update their decisions in static environments without individual state dynamics. In contrast, this work develops an evolutionary framework for dynamic games with explicit state evolution, which is necessary to model many applications. We introduce a mean field approximation of the finite-population game and establish approximation guarantees. Since state-of-the-art solution concepts for dynamic games lack an evolutionary interpretation, we propose a new concept - the Mixed Stationary Nash Equilibrium (MSNE) - which admits one. We characterize an equivalence between MSNE and the rest points of the proposed mean field evolutionary model and we give conditions for the evolutionary stability of MSNE.

eess.SY

Evolutionary Dynamics in Continuous-time Finite-state Mean Field Games -- Part II: Stability

We study a dynamic game with a large population of players who choose actions from a finite set in continuous time. Each player has a state in a finite state space that evolves stochastically with their actions. A player's reward depends not only on their own state and action but also on the distribution of states and actions across the population, capturing effects such as congestion in traffic networks. In Part I, we introduced an evolutionary model and a new solution concept - the mixed stationary Nash Equilibrium (MSNE) - which coincides with the rest points of the mean field evolutionary model under meaningful families of revision protocols. In this second part, we investigate the evolutionary stability of MSNE. We derive conditions on both the structure of the MSNE and the game's payoff map that ensure local and global stability under evolutionary dynamics. These results characterize when MSNE can robustly emerge and persist against strategic deviations, thereby providing insight into its long-term viability in large population dynamic games.

eess.SY

Evolutionary Dynamics in Continuous-time Finite-state Mean Field Games -- Part I: Equilibria

We study a dynamic game with a large population of players who choose actions from a finite set in continuous time. Each player has a state in a finite state space that evolves stochastically with their actions. A player's reward depends not only on their own state and action but also on the distribution of states and actions across the population, capturing effects such as congestion in traffic networks. While prior work in evolutionary game theory has primarily focused on static games without individual player state dynamics, we present the first comprehensive evolutionary analysis of such dynamic games. We propose an evolutionary model together with a mean field approximation of the finite-population game and establish strong approximation guarantees. We show that standard solution concepts for dynamic games lack an evolutionary interpretation, and we propose a new concept - the Mixed Stationary Nash Equilibrium (MSNE) - which admits one. We analyze the relationship between MSNE and the rest points of the mean field evolutionary model and study the evolutionary stability of MSNE.

eess.SY

Quantitative convergence of trained single layer neural networks to Gaussian processes

In this paper, we study the quantitative convergence of shallow neural networks trained via gradient descent to their associated Gaussian processes in the infinite-width limit. While previous work has established qualitative convergence under broad settings, precise, finite-width estimates remain limited, particularly during training. We provide explicit upper bounds on the quadratic Wasserstein distance between the network output and its Gaussian approximation at any training time $t \ge 0$, demonstrating polynomial decay with network width. Our results quantify how architectural parameters, such as width and input dimension, influence convergence, and how training dynamics affect the approximation error.

stat.ML

A multiscale analysis of mean-field transformers in the moderate interaction regime

In this paper, we study the evolution of tokens through the depth of encoder-only transformer models at inference time by modeling them as a system of particles interacting in a mean-field way and studying the corresponding dynamics. More specifically, we consider this problem in the moderate interaction regime, where the number $N$ of tokens is large and the inverse temperature parameter $\beta$ of the model scales together with $N$. In this regime, the dynamics of the system displays a multiscale behavior: a fast phase, where the token empirical measure collapses on a low-dimensional space, an intermediate phase, where the measure further collapses into clusters, and a slow one, where such clusters sequentially merge into a single one. We provide a rigorous characterization of the limiting dynamics in each of these phases and prove convergence in the above mentioned limit, exemplifying our results with some simulations.

cs.LG

Global Optimization via Softmin Energy Minimization

Global optimization, particularly for non-convex functions with multiple local minima, poses significant challenges for traditional gradient-based methods. While metaheuristic approaches offer empirical effectiveness, they often lack theoretical convergence guarantees and may disregard available gradient information. This paper introduces a novel gradient-based swarm particle optimization method designed to efficiently escape local minima and locate global optima. Our approach leverages a "Soft-min Energy" interacting function, $J_\beta(\mathbf{x})$, which provides a smooth, differentiable approximation of the minimum function value within a particle swarm. We define a stochastic gradient flow in the particle space, incorporating a Brownian motion term for exploration and a time-dependent parameter $\beta$ to control smoothness, similar to temperature annealing. We theoretically demonstrate that for strongly convex functions, our dynamics converges to a stationary point where at least one particle reaches the global minimum, with other particles exhibiting exploratory behavior. Furthermore, we show that our method facilitates faster transitions between local minima by reducing effective potential barriers with respect to Simulated Annealing. More specifically, we estimate the hitting times of unexplored potential wells for our model in the small noise regime and show that they compare favorably with the ones of overdamped Langevin. Numerical experiments on benchmark functions, including double wells and the Ackley function, validate our theoretical findings and demonstrate better performance over the well-known Simulated Annealing method in terms of escaping local minima and achieving faster convergence.

cs.LG

Noise-induced stabilization in a chemical reaction network without boundary effects

We present a chemical reaction network that is unstable under deterministic mass action kinetics, exhibiting finite-time blow-up of trajectories in the interior of the state space, but whose stochastic counterpart is positive recurrent. This provides an example of noise-induced stabilization of the model's dynamics arising due to noise perturbing transversally the divergent trajectories of the system that is independently of boundary effects. The proof is based on a careful decomposition of the state space and the construction of suitable Lyapunov functions in each region.

math.PR

Emergence of meta-stable clustering in mean-field transformer models

We model the evolution of tokens within a deep stack of Transformer layers as a continuous-time flow on the unit sphere, governed by a mean-field interacting particle system, building on the framework introduced in (Geshkovski et al., 2023). Studying the corresponding mean-field Partial Differential Equation (PDE), which can be interpreted as a Wasserstein gradient flow, in this paper we provide a mathematical investigation of the long-term behavior of this system, with a particular focus on the emergence and persistence of meta-stable phases and clustering phenomena, key elements in applications like next-token prediction. More specifically, we perform a perturbative analysis of the mean-field PDE around the iid uniform initialization and prove that, in the limit of large number of tokens, the model remains close to a meta-stable manifold of solutions with a given structure (e.g., periodicity). Further, the structure characterizing the meta-stable manifold is explicitly identified, as a function of the inverse temperature parameter of the model, by the index maximizing a certain rescaling of Gegenbauer polynomials.

cs.LG

Fair Artificial Currency Incentives in Repeated Weighted Congestion Games: Equity vs. Equality

When users access shared resources in a selfish manner, the resulting societal cost and perceived users' cost is often higher than what would result from a centrally coordinated optimal allocation. While several contributions in mechanism design manage to steer the aggregate users choices to the desired optimum by using monetary tolls, such approaches bear the inherent drawback of discriminating against users with a lower income. More recently, incentive schemes based on artificial currencies have been studied with the goal of achieving a system-optimal resource allocation that is also fair. In this resource-sharing context, this paper focuses on repeated weighted congestion game with two resources, where users contribute to the congestion to different extents that are captured by individual weights. First, we address the broad concept of fairness by providing a rigorous mathematical characterization of the distinct societal metrics of equity and equality, i.e., the concepts of providing equal outcomes and equal opportunities, respectively. Second, we devise weight-dependent and time-invariant optimal pricing policies to maximize equity and equality, and prove convergence of the aggregate user choices to the system-optimum. In our framework it is always possible to achieve system-optimal allocations with perfect equity, while the maximum equality that can be reached may not be perfect, which is also shown via numerical simulations.

cs.GT

Scalable and Calibrated Sampling for Bayesian Generalized Linear Mixed Model via Stochastic Gradient Markov Chain Monte Carlo

Generalized linear mixed models (GLMMs) are widely used for analyzing correlated data, particularly in large-scale biomedical and social science applications. Scalable Bayesian inference for GLMMs is challenging due to an intractable marginal likelihood and a high computational cost incurred by conventional Markov chain Monte Carlo (MCMC) methods. We develop a stochastic gradient MCMC (SGMCMC) algorithm tailored to GLMMs that enables accurate posterior inference in the large-sample regime. Our approach uses Fisher's identity to construct a (biased) Monte Carlo estimator of the gradient of the marginal log-likelihood, making SGMCMC feasible when direct gradient computation is impossible. We analyze the additional variability, introduced by both data subsampling and gradient approximation, to derive a post-hoc covariance correction that yields properly calibrated posterior uncertainty. We show through simulated studies that the proposed method provides accurate posterior means and variances in settings with a large number of groups, outperforming existing approaches, including control variate methods. We further demonstrate the method's practical utility in an analysis of electronic health records data, where accounting for variance inflation materially changes scientific conclusions.

stat.CO

Random Splitting of Point Vortex Flows

We consider a stochastic version of the point vortex system, in which the fluid velocity advects single vortices intermittently for small random times. Such system converges to the deterministic point vortex dynamics as the rate at which single components of the vector field are randomly switched diverges, and therefore it provides an alternative discretization of 2D Euler equations. The random vortex system we introduce preserves microcanonical statistical ensembles of the point vortex system, hence constituting a simpler alternative to the latter in the statistical mechanics approach to 2D turbulence.

math.PR

Global Optimality of Elman-type RNN in the Mean-Field Regime

We analyze Elman-type Recurrent Reural Networks (RNNs) and their training in the mean-field regime. Specifically, we show convergence of gradient descent training dynamics of the RNN to the corresponding mean-field formulation in the large width limit. We also show that the fixed points of the limiting infinite-width dynamics are globally optimal, under some assumptions on the initialization of the weights. Our results establish optimality for feature-learning with wide RNNs in the mean-field regime

stat.ML

Random Splitting of Fluid Models: Positive Lyapunov Exponents

In this paper we give sufficient conditions for random splitting systems to have a positive top Lyapunov exponent. We verify these conditions for random splittings of two fluid models: the conservative Lorenz-96 equations and Galerkin approximations of the 2D Euler equations on the torus. In doing so, we highlight particular structures in these equations such as shearing. Since a positive top Lyapunov exponent is an indicator of chaos which in turn is a feature of turbulence, our results show these randomly split fluid models have important characteristics of turbulent flow.

math.DS