SearcharxivSearch

arXiv subjects

Louis-Pierre Chaintron

Publications and source records attributed to Louis-Pierre Chaintron.

14 recordsLinked to original sources

A Dacorogna-Moser construction of transport maps on $\mathbb{R}^d$ with application to geodesics on the space of couplings

A seminal work by Dacorogna and Moser introduced a way of constructing regular transport maps from a probability distribution on a bounded domain to another one. In this work, we extend this construction to the whole $\mathbb{R}^d$ for strictly asymptotically log-concave measures, a wide class of distributions that encompasses Lipschitz-perturbations of log-concave measures. We then leverage this construction to study geodesics in the space of probability measures on a product set with imposed marginal laws (couplings), for which we derive optimality conditions, answering an open question in a recent work by Conforti, Lacker and Pal. Taking inspiration from Brenier's variational model for incompressible fluids and its regularization, we further introduce an entropic regularization of the geodesic problem, which can be seen as the Schr\"odinger bridge problem on the space of couplings, for which we also derive optimality conditions. We eventually study convergence of minimizers as the regularization vanishes, and we prove convergence of the related Lagrange multipliers. Our approach involves proving uniform-in-time global regularity estimates on elliptic and parabolic equations on $\mathbb{R}^d$, by exploiting the structure of asymptotically log-concave measures using the probabilistic notion of reflection coupling.

math.AP

ResNets of All Shapes and Sizes: Convergence of Training Dynamics in the Large-scale Limit

We establish convergence of the training dynamics of residual neural networks (ResNets) to their joint infinite depth L, hidden width M, and embedding dimension D limit. Specifically, we consider ResNets with two-layer perceptron blocks in the maximal local feature update (MLU) regime and prove that, after a bounded number of training steps, the error between the ResNet and its large-scale limit is O(1/L + sqrt(D/(L M)) + 1/sqrt(D)). This error rate is empirically tight when measured in embedding space. For a budget of P = Theta(L M D) parameters, this yields a convergence rate O(P^(-1/6)) for the scalings of (L, M, D) that minimize the bound. Our analysis exploits in an essential way the depth-two structure of residual blocks and applies formally to a broad class of state-of-the-art architectures, including Transformers with bounded key-query dimension. From a technical viewpoint, this work completes the program initiated in the companion paper [Chi25] where it is proved that for a fixed embedding dimension D, the training dynamics converges to a Mean ODE dynamics at rate O(1/L + sqrt(D)/sqrt(L M)). Here, we study the large-D limit of this Mean ODE model and establish convergence at rate O(1/sqrt(D)), yielding the above bound by a triangle inequality. To handle the rich probabilistic structure of the limit dynamics and obtain one of the first rigorous quantitative convergence for a DMFT-type limit, we combine the cavity method with propagation of chaos arguments at a functional level on so-called skeleton maps, which express the weight updates as functions of CLT-type sums from the past.

stat.ML

Mean-field limits \`a la Tanaka and large deviations for particle systems with network interactions

This article proposes a unified framework to study non-exchangeable mean-field particle systems with some general interaction mechanisms. The starting point is a fixed-point formulation of particle systems originally due to Tanaka that allows us to prove mean-field limit and large deviation results in an abstract setting. While it has been recently shown that such formulation encompasses a large class of exchangeable particle systems, we propose here a setting for the non-exchangeable case, including the case of adaptive interaction networks. We introduce sufficient conditions on the network structure that imply the mean-field limit and a new large deviations principle for the interaction measure. Finally, we formally highlight important models for which it is possible to derive a closed PDE characterization of the limit.

math.PR

Geodesic convexity and strengthened functional inequalities in submanifolds of Wasserstein space

We study the geodesic convexity of various energy and entropy functionals restricted to (non-geodesically convex) submanifolds of Wasserstein spaces with their induced geometry. We prove a variety of convexity results by means of a simple general principle, which holds in the metric space setting, and which crucially requires no knowledge of the structure of geodesics in the submanifold: If the EVI gradient flow of a functional exists and leaves the submanifold invariant, then the restriction of the functional to the submanifold is geodesically convex. This leads to short new proofs of several known results, such as one of Carlen and Gangbo on strong convexity of entropy on sphere-like submanifolds, and several new results, such as the $λ$-convexity of entropy on the space of couplings of $λ$-log-concave marginals. Along the way, we develop sufficient conditions for existence of geodesics in Wasserstein submanifolds. Submanifold convexity results lead systematically to improvements of Talagrand and HWI inequalities which we speculate to be closely related to concentration of measure estimates for conditioned empirical measures, and we prove one rigorous result in this direction in the Carlen-Gangbo setting.

math.AP

Propagation of weak log-concavity along generalised heat flows via Hamilton-Jacobi equations

A well-known consequence of the Pr{é}kopa-Leindler inequality is the preservation of logconcavity by the heat semigroup. Unfortunately, this property does not hold for more general semigroups. In this paper, we exhibit a slightly weaker notion of log-concavity that can be propagated along generalised heat semigroups. As a consequence, we obtain logsemiconcavity properties for the ground state of Schr{ö}dinger operators for non-convex potentials, as well as propagation of functional inequalities along generalised heat flows. We then investigate the preservation of weak log-concavity by conditioning and marginalisation, following the seminal works of Brascamp and Lieb. To our knowledge, our results are the first of this type in non log-concave settings. We eventually study generation of log-concavity by parabolic regularisation and prove novel two-sided log-Hessian estimates for the fundamental solution of parabolic equations with unbounded coefficients, which can be made uniform in time. These properties are obtained as a consequence of new propagation of weak convexity results for quadratic Hamilton-Jacobi-Bellman (HJB) equations. The proofs rely on a stochastic control interpretation combined with a second order analysis of reflection coupling along HJB characteristics.

math.AP

Optimal rate of convergence in the vanishing viscosity for uniformly convex Hamilton-Jacobi equations

The purpose of this note is to provide an optimal rate of convergence in the vanishing viscosity regime for first-order Hamilton-Jacobi equations with uniformly convex Hamiltonian. We prove that for a globally Lipschitz-continuous and semiconcave terminal condition the rate is of order O($ε$log$ε$), and we provide an example to show that this rate cannot be sharpened. This improves on the previously known rate of convergence O($\sqrt$$ε$), which was widely believed to be optimal. Our proof combines techniques involving regularisation by sup-convolution with entropy estimates for the flow of a suitable version of the adjoint linearized equation. The key technical point is an integrated estimate of the Laplacian of the solution against this flow. Moreover, we exploit the semiconcavity generated by the equation to handle less regular data in the quadratic case.

math.AP

Optimal rate of convergence in the vanishing viscosity for quadratic Hamilton-Jacobi equations

The purpose of this note is to provide an optimal rate of convergence in the vanishing viscosity regime for first-order Hamilton-Jacobi equations with purely quadratic Hamiltonian. We show that for a globally Lipschitz-continuous terminal condition the rate is of order O($ε$ log $ε$), and we provide an example to show that this rate cannot be sharpened. This improves on the previously known rate of convergence O( $\sqrt$ $ε$), which was widely believed to be optimal. Our proof combines techniques involving regularization by sup-convolution with entropy estimates for the flow of a suitable version of the adjoint linearized equation. The key technical point is an integrated estimate of the Laplacian of the solution against this flow. Moreover, we exploit the semiconcavity generated by the equation.

math.AP

Constrained non-linear estimation and links with stochastic filtering

This article studies the problem of estimating the state variable of non-smooth subdifferential dynamics constrained in a bounded convex domain given some real-time observation. On the one hand, we show that the value function of the estimation problem is a viscosity solution of a Hamilton Jacobi Bellman equation whose sub and super solutions have different Neumann type boundary conditions. This intricacy arises from the non-reversibility in time of the non-smooth dynamics, and hinders the derivation of a comparison principle and the uniqueness of the solution in general. Nonetheless, we identify conditions on the drift (including zero drift) coefficient in the non-smooth dynamics that make such a derivation possible. On the other hand, we show in a general situation that the value function appears in the small noise limit of the corresponding stochastic filtering problem by establishing a large deviation result. We also give quantitative approximation results when replacing the non-smooth dynamics with a smooth penalised one.

math.OC

Regularity and stability for the Gibbs conditioning principle on path space via McKean-Vlasov control

We consider a system of diffusion processes interacting through their empirical distribution. Assuming that the empirical average of a given observable can be observed at any time, we derive regularity and quantitative stability results for the optimal solutions in the associated version of the Gibbs conditioning principle. The proofs rely on the analysis of a McKean-Vlasov control problem with distributional constraints. Some new estimates are derived for Hamilton-Jacobi-Bellman equations and the Hessian of the log-density of diffusion processes, which are of independent interest.

math.OC

Gibbs principle with infinitely many constraints: optimality conditions and stability

We extend the Gibbs conditioning principle to an abstract setting combining infinitely many linear equality constraints and non-linear inequality constraints, which need not be convex. A conditional large large deviation principle (LDP) is proved in a Wassersteintype topology, and optimality conditions are written in this abstract setting. This setting encompasses versions of the Schr{ö}dinger bridge problem with marginal non-linear inequality constraints at every time. In the case of convex constraints, stability results for perturbations both in the constraints and the reference measure are proved. We then specify our results when the reference measure is the path-law of a continuous diffusion process, whose law is constrained at each time. We obtain a complete description of the constrained process through an atypical mean-field PDE system involving a Lagrange multiplier.

math.FA

Quasi-continuity method for mean-field diffusions: large deviations and central limit theorem

A pathwise large deviation principle in the Wasserstein topology and a pathwise central limit theorem are proved for the empirical measure of a mean-field system of interacting diffusions. The coefficients are path-dependent. The framework allows for degenerate diffusion matrices, which may depend on the empirical measure, including mean-field kinetic processes. The main tool is an extension of Tanaka's pathwise construction to non-constant diffusion matrices. This can be seen as a mean-field analogous of Azencott's quasi-continuity method for the Freidlin-Wentzell theory. As a by-product, uniform-in-time-step fluctuation and large deviation estimates are proved for a discrete-time version of the meanfield system. Uniform-in-time-step convergence is also proved for the value function of some mean-field control problems with quadratic cost.

math.PR

Existence and global Lipschitz estimates for unbounded classical solutions of a Hamilton-Jacobi equation

The purpose of this article is to prove existence, uniqueness and uniform gradient estimates for unbounded classical solutions of a Hamilton-Jacobi-Bellman equation. Such an equation naturally arises in stochastic control problems. Contrary to the classical literature which handles the case of bounded regular coefficients, we only impose Lipschitz regularity conditions, allowing for a linear growth of coefficients. This latter regularity assumption is natural in a probabilistic setting. In principle, this assumption is compatible with global Lipschitz regularity for the solution. However, to the best of our knowledge, this useful result had not been established before. The proof that we provide relies on classical methods from the viscosity solution theory, combining the Ishii-Lions method [IL90] for uniformly elliptic equations with ideas from the weak Bernstein method [Bar91].

math.AP

Propagation of chaos: a review of models, methods and applications. I. Models and methods

The notion of propagation of chaos for large systems of interacting particles originates in statistical physics and has recently become a central notion in many areas of applied mathematics. The present review describes old and new methods as well as several important results in the field. The models considered include the McKean-Vlasov diffusion, the mean-field jump models and the Boltzmann models. The first part of this review is an introduction to modelling aspects of stochastic particle systems and to the notion of propagation of chaos. The second part presents concrete applications and a more detailed study of some of the important models in the field.

math.PR

Propagation of chaos: a review of models, methods and applications. II. Applications

The notion of propagation of chaos for large systems of interacting particles originates in statistical physics and has recently become a central notion in many areas of applied mathematics. The present review describes old and new methods as well as several important results in the field. The models considered include the McKean-Vlasov diffusion, the mean-field jump models and the Boltzmann models. The first part of this review is an introduction to modelling aspects of stochastic particle systems and to the notion of propagation of chaos. The second part presents concrete applications and a more detailed study of some of the important models in the field.

math.PR