SearcharxivSearch

arXiv subjects

Qinxin Yan

Publications and source records attributed to Qinxin Yan.

10 recordsLinked to original sources

Finite-player Optimal Stopping Games: Randomization, $\alpha$-potentiality, and Learning

Finite-player nonzero-sum optimal stopping games typically lead to coupled equilibrium systems whose complexity grows rapidly with the number of players. We introduce an independently randomized formulation in which each stopping rule is represented by an adapted, nondecreasing cumulative stopping process. The canonical embedding preserves pure-profile payoffs, and a pure profile is a Nash equilibrium of the original game if and only if its embedding is a Nash equilibrium of the randomized game. We adopt the $\alpha$-potential approach to construct an $\alpha_N$-potential function, with the error $\alpha_N=O(N^{-1})$ under weak-interaction. We also identify an exact-potential subclass with a closed-form threshold equilibrium. For local stopped-status interactions, randomized payoffs admit a local stopped-mass representation, and potential maximization can be formulated as a multidimensional singular-control problem with local gradient constraints and a nonlocal condition for finite jumps. Under suitable regularity assumptions, we study the associated Hamilton-Jacobi-Bellman quasi-variational inequality and its regularity properties. For unknown model coefficients, we propose a bounded-intensity Potential-CT-DDPG learning algorithm. Numerical experiments closely match the analytical benchmark and yield estimated best-response improvements consistent with $N^{-1}$ scaling.

math.OC

Optimal Loss Allocation in a Mean-Field Model of Systemic Risk

We study a systemic-risk control problem in which a central planner allocates losses generated by bank defaults across the surviving institutions. Banks are modeled through their distances to default, evolving as absorbed Brownian motions with downward jumps induced by redistributed default losses. Unlike bailout models, the planner cannot inject external capital or reduce the aggregate loss, and the only admissible intervention is to decide how each endogenous loss is assigned among solvent banks. The objective is to maximize terminal system health, including survival mass as a leading special case and, more generally, increasing concave welfare functionals of the terminal distribution. Our main result identifies an optimal allocation rule with a simple economic interpretation: losses should be concentrated on the currently healthiest institutions. In discrete time, this rule takes the form of a cutoff or taxing-the-richest policy, which reduces banks above an endogenous threshold down to that threshold while leaving weaker banks untouched. We prove convergence of the time-discretized mean-field control problem as the allocation time step tends to zero and characterize the limiting problem as a singular mean-field control problem. The optimally controlled law is described by a reflected free-boundary formulation, in which the cutoff becomes the moving upper edge of the support, and the associated value function satisfies a Hamilton-Jacobi equation on Wasserstein space. Finally, we formulate the corresponding finite-particle control problem and show, under suitable assumptions, that the cutoff-controlled particle system converges to the continuous-time mean-field model. This provides a finite-system foundation for the optimal mean-field loss-allocation rule.

math.OC

Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

This paper develops a model-free framework for continuous-time mean-field control when the population evolves according to unknown controlled McKean--Vlasov dynamics and only discrete-time transition data are available. Model-based mean-field control requires the continuous-time drift and diffusion coefficients, which are not directly observed from fixed-step transitions, while a direct reduction to a discrete-time Bellman equation loses the continuous-time generator structure. To bridge these two viewpoints, we introduce a Mean-Field-PhiBE (MF-PhiBE), which incorporates discrete-time transition information into a continuous-time PDE on the Wasserstein space. The MF-PhiBE replaces the unknown infinitesimal drift and covariance in the policy-evaluation equation by one-step estimators computed from data, while preserving the generator structure of the McKean-Vlasov HJB equation. We also derive a policy-gradient theorem for entropy-regularized randomized feedback policies, expressing the actor direction through an action-wise infinitesimal advantage and the score of the policy. Combining these two ingredients yields a model-free actor-critic method. We prove a first-order consistency estimate showing that the value induced by an optimal MF-PhiBE policy approximates the optimal continuous-time value as the observation time step vanishes. For entropy-regularized LQR, we establish first-order policy convergence and second-order value convergence; under suitable conditions, the population-averaged feedback means coincide exactly. Numerical experiments on an LQR benchmark and a crowd-aversion problem illustrate the proposed framework.

math.OC

Policy Gradient for Continuous-Time Mean-Field Control

This paper develops a policy gradient method for entropy-regularized mean-field control in the discounted infinite-horizon setting. We consider randomized feedback policies and a coupled representative-particle/population system, in which the representative state evolves jointly with a population law governed by a McKean--Vlasov equation. The resulting value function is therefore defined on the product space $\mathbb R^d \times \mathcal P_2(\mathbb R^d)$. A key distinction from existing policy gradient methods for mean-field control is that, after computing the value function under a fixed policy, our approach does not require solving an additional equation to obtain the policy gradient. Instead, we derive an explicit policy gradient formula directly in terms of the value function. The formulation is based on an instantaneous advantage function, which quantifies the gain of taking a given action relative to the current randomized policy. We establish a G\^ateaux policy-gradient formula, which gives the first-order variation of the objective along arbitrary policy perturbations, and then derive the corresponding ascent direction under finite-dimensional policy parametrization. The resulting formula leads to a model-based actor--critic scheme. The critic is obtained by solving the associated linear stationary Hamilton--Jacobi--Bellman equation for the value function, using cylindrical functions to represent dependence on the population law. The actor is then updated according to the derived policy-gradient formula. We further analyze the well-posedness of the PDE in a polynomial-growth function class. Finally, we illustrate the proposed method through numerical experiments on an LQR model and a crowd-motion problem.

math.OC

Implicit Regularization of Large Neural Networks via Mean-Field Formulation

We propose a mathematical framework to explain implicit regularization from early stopping during the training of overparametrized neural networks. In the mean-field limit, the parameter distribution evolves according to a gradient flow on the space of probability measures. We show that these dynamics admit an equivalent McKean-Vlasov stochastic control formulation through the corresponding Hamilton-Jacobi-Bellman (HJB) equation. The control viewpoint yields a Dynamic Programming Principle (DPP), which we use to define a new metric on probability measures. This metric can be viewed as a mean-field generalization of the control representation of the Wasserstein-2 distance, and it naturally appears as a regularization term selected by early stopping. We further obtain non-asymptotic bounds describing how the induced regularization depends on the stopping time.

math.OC

Iterative Schemes for Markov Perfect Equilibria

We study Markov perfect equilibria in continuous-time dynamic games with finitely many symmetric players. The corresponding Nash system reduces to the Nash-Lasry-Lions equation for the common value function, also known as the master equation in the mean-field setting. In the finite-state space problems we consider, this equation becomes a nonlinear ordinary differential equation admitting a unique classical solution. Leveraging this uniqueness, we prove the convergence of both Picard and weighted Picard iterations, yielding efficient computational methods. Numerical experiments confirm the effectiveness of algorithms based on this approach.

math.OC

Particle Systems with Local Interactions via Hitting Times and Cascades on Graphs

We introduce a family of particle systems on sparse graphs where local interactions occur via hitting times, providing a dynamic and tractable model for default cascades in large sparsely-connected financial networks. Building on the framework of Lacker, Ramanan and Wu (2023), we extend convergence theory to systems with singular interactions, capturing the abrupt and discontinuous nature of systemic events. We establish conditions for well-posedness through a minimality principle and connect fragility to dynamic percolation thresholds. Our analysis demonstrates continuity of the joint law of defaults with respect to local graph convergence, establishes convergence of empirical distributions, and characterizes the default time distribution in tree-like networks. This framework offers a rigorous and flexible foundation for modeling systemic risk in evolving financial systems, featuring continuous-time dynamics, heterogeneous and local interactions, and instantaneous default cascades.

math.PR

Learning algorithms for mean field optimal control

We analyze an algorithm to numerically solve the mean-field optimal control problems by approximating the optimal feedback controls using neural networks with problem specific architectures. We approximate the model by an $N$-particle system and leverage the exchangeability of the particles to obtain substantial computational efficiency. In addition to several numerical examples, a convergence analysis is provided. We also developed a universal approximation theorem on Wasserstein spaces.

math.OC

Viscosity Solutions of the Eikonal Equation on the Wasserstein Space

Dynamic programming equations for mean field control problems with a separable structure are Eikonal equations on the Wasserstein space. Standard differentiation using linear derivatives yield a direct extension of the classical viscosity theory. We use Fourier representation of the Sobolev norms on the space of measures, together with the standard techniques from the finite dimensional theory to obtain a comparison result among sub and super solutions that are Lipschitz continuous in the Wasserstein-one norm which is directly verified for the value function. A key technical result provides a point-wise upper bound for the derivative of the linear derivative of a function in terms of its Lipschitz norm.

math.OC

Viscosity Solutions for McKean-Vlasov Control on a torus

An optimal control problem in the space of probability measures, and the viscosity solutions of the corresponding dynamic programming equations defined using the intrinsic linear derivative are studied. The value function is shown to be Lipschitz continuous with respect to a novel smooth Fourier Wasserstein metric. A comparison result between the Lipschitz viscosity sub and super solutions of the dynamic programming equation is proved using this metric, characterizing the value function as the unique Lipschitz viscosity solution.

math.OC