SearcharxivSearch

arXiv subjects

Pierre Monmarché

Publications and source records attributed to Pierre Monmarché.

At least 19 recordsLinked to original sources

Equivalence between $N$-particle log-Sobolev inequalities and non-linear Łojasiewicz inequalities for a class of mean field systems

The convergence rate of a free energy Wasserstein gradient flow is quantified by its so-called Polyak-Lojasiewicz (PL) constant $λ$, which relates the objective function to its dissipation along the flow. Such a flow is the mean-field limit as $N$ goes to infinity of a system of $N$ interacting particles, whose convergence rate in relative entropy towards its Gibbs measure is quantified by its log-Sobolev constant $λ_N$. Different behaviours of $λ_N$ as $N$ goes to infinity thus describe drastically different phenomena, such as fast relaxation or metastability, with many models undergoing phase transitions between these regimes, depending typically on temperature. Under fairly general conditions, a uniform-in-$N$ log-Sobolev constant (i.e. fast exponential convergence for the particle system) is known to induce a positive PL constant (i.e. exponential convergence for the mean-field flow). A conjecture was stated by Delgadino, Gvalani, Pavliotis and Smith according to which the converse implication was true ($λ>0$ implies $\liminf λ_N >0$), even with $\lim λ_N =λ$. First, we will prove this converse implication, although without the equality $\lim λ_N = λ$, for a general class of mean-field models. Second, we also notice that this implication fails if the free energy minimiser is not unique, and provide an explicit counter-example. Third, we also consider the same question of relating $N$-particle and mean-field inequalities in the context of more general Lojasiewicz inequalities, which correspond to polynomial (instead of exponential) convergence rates, and can describe the situation exactly at a phase transition.

math.FA

Local exponential stability of mean-field Langevin descent-ascent and associated particle system

We study the mean-field Langevin descent-ascent (MFL-DA), a coupled optimization dynamics on the space of probability measures for entropically regularized two-player zero-sum games, together with its associated interacting particle system. For general nonconvex-nonconcave payoffs, Wang and Chizat (COLT 2024) asked whether the original single-timescale MFL-DA converges to the mixed Nash equilibrium and, if so, at what rate. We prove a local affirmative answer in Wasserstein space: if the initial datum is sufficiently close to the mixed Nash equilibrium, then the mean-field dynamics converges to it exponentially fast at a quantitative rate. We further show that the finite-$N$ particle system inherits this stability up to times exponential in $N$, with an $N$-independent exponential rate modulo a finite-particle error floor. Combined with the recent counterexample of Mourrat and Pillaud-Vivien for MFL-DA, which shows that global convergence cannot hold in general, our theorem completes the positive local counterpart of the Wang-Chizat question: the mixed Nash equilibrium has a robust basin of attraction, stable under both the mean-field flow and its finite-particle approximation.

cs.LG

On the entropic convergence for piecewise deterministic samplers: speedup and obstruction

For piecewise deterministic samplers such as Randomized Hamiltonian Monte Carlo (RHMC), Bouncy Particle Sampler (BPS) or Zig-Zag Process (ZZP), long-time exponential convergence rates have been established in previous works using Harris or $L^2$ hypocoercivity approaches. In particular, in the $L^2$ framework, a so-called \emph{diffusive-to-ballistic} speedup was known for log-concave targets, according to which the convergence rates of these samplers, with suitable parameters, are quadratically improved with respect to the standard overdamped Langevin diffusion process. A recent work by Jianfeng Lu showed that this speedup also holds for the kinetic Langevin diffusion process when the convergence is stated in terms of relative entropy, raising the question whether this also holds for piecewise deterministic samplers. The present work provides a positive and a negative answer to this: first, we show that the speedup holds in entropy for RHMC; second, we show that for BPS or ZZS, even for a standard Gaussian target, a similar result cannot hold, and even that exponential convergence (at any rate) in entropy fails.

math.PR

Quantum Circuits for the Metropolis-Hastings Algorithm

Szegedy's quantization of a reversible Markov chain provides a quantum walk whose spectral gap is quadratically larger than that of the classical walk. Quantum computers are therefore expected to provide a speedup of Metropolis-Hastings (MH) simulations. Existing generic methods to implement the quantum walk require coherently computing the transition probabilities of the underlying Markov kernel. However, reversible computing methods require a number of qubits that scales with the complexity of the computation. This overhead is undesirable in near-term fault-tolerant quantum computing, where few logical qubits are available. In this work, we present a Szegedy quantum walk construction which follows the classical proposal-acceptance logic, and does not require further reversible computing methods. We also compare this construction with an alternative to Szegedy's approach which also provides a quadratic gap amplification. Since each step of the quantum walks uses a constant number of proposal and acceptance steps, we expect the end-to-end quadratic speedup to hold for MH Markov Chain Monte-Carlo simulations.

quant-ph

Nesterov acceleration for the Wasserstein minimization of displacement-convex free energies

We show that the mean-field underdamped Langevin process (associated to the non-linear Vlasov-Fokker-Planck equation) achieves a Nesterov acceleration with respect to the Wasserstein gradient flow of a displacement-convex free energy, in the sense that it converges at a rate of order given by the square-root of the Polyak-Łojasiewicz constant of the free energy (which is the optimal convergence rate for the corresponding gradient flow). This result has been made possible by the recent breakthrough [42] by Jianfeng Lu, which establishes such a \emph{diffusive-to-ballistic} improvement in term of entropy in the linear case.

math.AP

Long-time $L^p$ Wasserstein contraction for diffusion processes without global dissipativity

The fact that a Markov diffusion semi-group on $\mathbb R^d$ contracts the $L^p$ Wasserstein distance, which has been extensively used to establish uniform-in-time stability estimates (e.g. with respect to numerical discretization errors), is a well-studied question in the case where the distances are in fact deterministically contracted by the drift (global dissipativity condition) or in the case $p=1$ (with reflection couplings). This work focuses on the non-globally dissipative case with $p>1$. This situation was previously considered in \cite{MonmarcheBruit}, but only for elliptic processes, and with a restriction on the diffusivity coefficient (which had to be large enough). Here, we extend this analysis to non-elliptic processes and provide sharper conditions to get contractions along synchronous coupling, including negative results, lower bounds and a characterization (at least in dimension 1) in terms of the maximal eigenvalue of a Feynman-Kac operator.

math.PR

General-purpose post-sampling reweighting method for multimodal target measures

When sampling multi-modal probability distributions, correctly estimating the relative probability of each mode, even when the modes have been discovered and locally sampled, remains challenging. We test a simple reweighting scheme designed for this situation, which consists in minimizing (in terms of weights) the Kullback-Leibler divergence of a weighted (regularized) empirical distribution of the samples with respect to the target measure.

math.ST

Convergence rates for an Adaptive Biasing Potential scheme from a Wasserstein optimization perspective

Free-energy-based adaptive biasing methods, such as Metadynamics, the Adaptive Biasing Force (ABF) and their variants, are enhanced sampling algorithms widely used in molecular simulations. Although their efficiency has been empirically acknowledged for decades, providing theoretical insights via a quantitative convergence analysis is a difficult problem, in particular for the kinetic Langevin diffusion, which is non-reversible and hypocoercive. We obtain the first exponential convergence result for such a process, in an idealized setting where the dynamics can be associated with a mean-field non-linear flow on the space of probability measures. A key of the analysis is the interpretation of the (idealized) algorithm as the gradient descent of a suitable functional over the space of probability distributions.

math.PR

A Lyapunov-tamed Euler method for singular SDEs

Many applications, such as systems of interacting particles in physics, require the simulation of diffusion processes with singular coefficients. Standard Euler schemes are then not convergent, and theoretical guarantees in this situation are scarce. In this work we introduce a Lyapunov-tamed Euler scheme, for drift coefficients for which the weak derivative is dominated by a function that obeys a certain generic Lyapunov-type condition. This allows for a range of coefficients that explode to infinity on a bounded set. We establish that, in terms of Lp-strong error, the Lyapunov-tamed scheme is consistent and moreover achieves the same order of convergence as the standard Euler scheme for Lipschitz coefficients. The general result is applied to systems of mean-field particles with singular repulsive interaction in 1D, yielding an error bound with polynomial dependency in the number of particles.

math.PR

Piecewise deterministic sampling with splitting schemes

We introduce Markov chain Monte Carlo (MCMC) algorithms based on numerical approximations of piecewise-deterministic Markov processes obtained with the framework of splitting schemes. We present unadjusted as well as adjusted algorithms, for which the asymptotic bias due to the discretisation error is removed applying a non-reversible Metropolis-Hastings filter. In a general framework we demonstrate that the unadjusted schemes have weak error of second order in the step size, while typically maintaining a computational cost of only one gradient evaluation of the negative log-target function per iteration. Focusing then on unadjusted schemes based on the Bouncy Particle and Zig-Zag samplers, we provide conditions ensuring geometric ergodicity and consider the expansion of the invariant measure in terms of the step size. We analyse the dependence of the leading term in this expansion on the refreshment rate and on the structure of the splitting scheme, giving a guideline on which structure is best. Finally, we illustrate promising results for our samplers with numerical experiments on a Bayesian imaging inverse problem and a system of interacting particles.

math.PR

Free energy Wasserstein gradient flow and their particle counterparts: toy model, (degenerate) PL inequalities and exit times

In finite dimension, the long-time and metastable behavior of a gradient flow perturbated by a small Brownian noise is well understood. A similar situation arises when a Wasserstein gradient flow over a space of probability measure is approximated by a system of mean-field interacting particles, but classical results do not apply in these infinite-dimensional settings. This work is concerned with the situation where the objective function of the optimization problem contains an entropic penalization, so that the particle system is a Langevin diffusion process. We consider a very simple class of models, for which the infinite-dimensional behavior is fully characterized by a finite-dimensional process. The goal is to have a flexible class of benchmarks to fix some objectives, conjectures and (counter-)examples for the general situation. Inspired by the systematic study of these toy models, one application is presented on the continuous Curie-Weiss model in a symmetric double-well potential. We show that, at the critical temperature, although the $N$-particle Gibbs measure does not satisfy a uniform-in-$N$ standard log-Sobolev inequality (the optimal constant growing like $\sqrt{N}$), it does satisfy a more general Lojasiewicz inequality uniformly in $N$, inducing uniform polynomial long-time convergence rates, propagation of chaos at stationarity and uniformly in time, and creation of chaos.

math.PR

Quantum Speedup for Nonreversible Markov Chains

Quantum algorithms can potentially solve a handful of problems more efficiently than their classical counterparts. In that context, it has been discussed that Markov chains problems could be solved significantly faster using quantum computing. Indeed, previous work suggests that quantum computers could accelerate sampling from the stationary distribution of reversible Markov chains. However, in practice, certain physical processes of interest are nonreversible in the probabilistic sense and reversible Markov chains can sometimes be replaced by more efficient nonreversible chains targeting the same stationary distribution. This study constructs Markov chain reversibilizations and develops quantum algorithmic techniques to accelerate nonreversible processes. Such an up-to-exponential quantum speedup goes beyond the predicted quadratic quantum acceleration for reversible chains and is likely to have a decisive impact on many applications ranging from statistics and machine learning to computational modeling in physics, chemistry, biology and finance.

quant-ph

Exponential Ergodicity in Relative Entropy and $L^2$-Wasserstein Distance for non-equilibrium partially dissipative Kinetic SDEs

In this paper, we derive exponential ergodicity in relative entropy for general kinetic SDEs under a partially dissipative condition. It covers non-equilibrium situations where the forces are not of gradient type and the invariant measure does not have an explicit density, extending previous results set in the equilibrium case. The key argument is to establish the hypercontractivity of the associated semigroup, which follows from its hyperboundedness and its $L^2$-exponential ergodicity. Moreover, we obtain exponential ergodicity in the $L^2$-Wasserstein distance by combining Talagrand's inequality with a log-Harnack inequality. These results are further extended to the McKean-Vlasov setting and to the associated mean-field interacting particle systems, with convergence rates that are uniform in the number of particles in the latter case, under small nonlinear perturbations.

math.PR

Local convergence rates for Wasserstein gradient flows and McKean-Vlasov equations with multiple stationary solutions

Non-linear versions of log-Sobolev inequalities, that link a free energy to its dissipation along the corresponding Wasserstein gradient flow (i.e. corresponds to Polyak-Lojasiewicz inequalities in this context), are known to provide global exponential long-time convergence to the free energy minimizers, and have been shown to hold in various contexts. However they cannot hold when the free energy admits critical points which are not global minimizers, which is for instance the case of the granular media equation in a double-well potential with quadratic attractive interaction at low temperature. This work addresses such cases, extending the general arguments when a log-Sobolev inequality only holds locally and, as an example, establishing such local inequalities for the granular media equation with quadratic interaction either in the one-dimensional symmetric double-well case or in higher dimension in the low temperature regime. The method provides quantitative convergence rates for initial conditions in a Wasserstein ball around the stationary solutions. The same analysis is carried out for the kinetic counterpart of the gradient flow, i.e. the corresponding Vlasov-Fokker-Planck equation. The local exponential convergence to stationary solutions for the mean-field equations, both elliptic and kinetic, is shown to induce for the corresponding particle systems a fast (i.e. uniform in the number or particles) decay of the particle system free energy toward the level of the non-linear limit.

math.AP

The velocity jump Langevin process and its splitting scheme: long time convergence and numerical accuracy

The Langevin dynamics is a diffusion process extensively used, in particular in molecular dynamics simulations, to sample Gibbs measures. Some alternatives based on (piecewise deterministic) kinetic velocity jump processes have gained interest over the last decade. One interest of the latter is the possibility to split forces (at the continuous-time level), reducing the numerical cost for sampling the trajectory. Motivated by this, a numerical scheme based on hybrid dynamics combining velocity jumps and Langevin diffusion, numerically more efficient than their classical Langevin counterparts, has been introduced for computational chemistry in [42]. The present work is devoted to the numerical analysis of this scheme. Our main results are, first, the exponential ergodicity of the continuous-time velocity jump Langevin process, second, a Talay-Tubaro expansion of the invariant measure of the numerical scheme on the torus, showing in particular that the scheme is of weak order 2 in the step-size and, third, a bound on the quadratic risk of the corresponding practical MCMC estimator (possibly with Richardson extrapolation). With respect to previous works on the Langevin diffusion, new difficulties arise from the jump operator, which is non-local.

math.NA

Stochastic moments dynamics: a flexible finite-dimensional random perturbation of Wasserstein gradient descent

For optimizing a non-convex function in finite dimension, a method is to add Brownian noise to a gradient descent, allowing for transitions between basins of attractions of different minimizers. To adapt this for optimization over a space of probability distributions requires a suitable noise. For this purpose, we introduce here a simple stochastic process where a number of moments of the distribution are following a chosen finite-dimensional diffusion process, generalizing some previous studies where the expectation of the measure is subject to a Brownian noise. The process may explode in finite time, for instance when trying to force the variance of a distribution to behave like a Brownian motion. We show, up to the possible explosion time, well-posedness and propagation of chaos for the system of mean-field interacting particles with common noise approximating the process.

math.PR

Convergence of Time-Averaged Mean Field Gradient Descent Dynamics for Continuous Multi-Player Zero-Sum Games

The approximation of mixed Nash equilibria (MNE) for zero-sum games with mean-field interacting players has recently raised much interest in machine learning. In this paper we propose a mean-field gradient descent dynamics for finding the MNE of zero-sum games involving $K$ players with $K\geq 2$. The evolution of the players' strategy distributions follows coupled mean-field gradient descent flows with momentum, incorporating an exponentially discounted time-averaging of gradients. First, in the case of a fixed entropic regularization, we prove an exponential convergence rate for the mean-field dynamics to the mixed Nash equilibrium with respect to the total variation metric. This improves a previous polynomial convergence rate for a similar time-averaged dynamics with different averaging factors. Moreover, unlike previous two-scale approaches for finding the MNE, our approach treats all player types on the same time scale. We also show that with a suitable choice of decreasing temperature, a simulated annealing version of the mean-field dynamics converges to an MNE of the initial unregularized problem.

math.OC

Logarithmic Sobolev inequalities for non-equilibrium steady states

We consider two methods to establish log-Sobolev inequalities for the invariant measure of a diffusion process when its density is not explicit and the curvature is not positive everywhere. In the first approach, based on the Holley-Stroock and Aida-Shigekawa perturbation arguments [J. Stat. Phys., 46(5-6):1159-1194, 1987, J. Funct. Anal., 126(2):448-475, 1994], the control on the (non-explicit) perturbation is obtained by stochastic control methods, following the comparison technique introduced by Conforti [Ann. Appl. Probab., 33(6A):4608-4644, 2023]. The second method combines the Wasserstein-2 contraction method, used in [Ann. Henri Lebesgue, 6:941-973, 2023] to prove a Poincaré inequality in some non-equilibrium cases, with Wang's hypercontractivity results [Potential Anal., 53(3):1123-1144, 2020].

math.PR