Searcharxiv⌕ Search

arXiv subjects

Mateusz B. Majka

Publications and source records attributed to Mateusz B. Majka.

At least 19 recordsLinked to original sources

On the computational cost of Stochastic Gradient Langevin Dynamics

Stochastic Gradient Langevin Dynamics (SGLD) reduces the cost of Langevin-based sampling by replacing full-dataset drift evaluations with mini-batch approximations, but the resulting subsampling error may offset this computational saving. We study this trade-off for stochastic differential equations with finite-sum drifts and compare the computational cost of SGLD with that of the Euler-Maruyama (EM) method. For a prescribed mean-square accuracy $\varepsilon^2$, we derive complexity estimates that explicitly track the dependence on the dataset size $m$, mini-batch size $s$, and accuracy parameter $\varepsilon$. The resulting comparison reveals distinct parameter regimes in which either method is preferable. In particular, EM can have lower leading-order cost only in a small-data, aggressive-subsampling regime, whereas SGLD is favoured over most of the remaining parameter space. In the practically relevant regime $s \ll m$, the transition between the two methods occurs at the scale $m \asymp \varepsilon^{-1}$. We complement the theoretical analysis with numerical experiments based on a Gaussian Bayesian inference model, which examine the predicted cost regimes together with the underlying discretisation error and variance estimates.

math.NA↗

Linear convergence of proximal descent schemes on the Wasserstein space

We investigate proximal descent methods, inspired by the minimizing movement scheme introduced by Jordan, Kinderlehrer and Otto, for optimizing entropy-regularized functionals on the Wasserstein space. We establish linear convergence under flat convexity assumptions, thereby relaxing the common reliance on geodesic convexity. Our analysis circumvents the need for discrete-time adaptations of the Evolution Variational Inequality (EVI). Instead, we leverage a uniform logarithmic Sobolev inequality (LSI) and the entropy ``sandwich" lemma, extending the analysis from arXiv:2201.10469 and arXiv:2202.01009. The major challenge in the proof via LSI is to show that the relative Fisher information is well-defined at every step of the scheme. Since the relative entropy is not Wasserstein differentiable, we prove that along the scheme the iterates belong to a certain class of Sobolev regularity, and hence the relative entropy has a unique Wasserstein sub-gradient, and that the relative Fisher information is indeed finite.

math.OC↗

Mirror Descent-Ascent for mean-field min-max problems

We study two variants of the mirror descent-ascent (MDA) algorithm for solving min-max problems on the space of measures: simultaneous and alternating. We work under assumptions of convexity-concavity and relative smoothness of the payoff function with respect to a suitable Bregman divergence, defined on the space of measures via flat derivatives. We establish non-asymptotic convergence rates to mixed Nash equilibria, measured in the Nikaidô-Isoda error, proving an $\mathcal{O}(N^{-1/2})$ rate for simultaneous MDA and an improved $\mathcal{O}(N^{-2/3})$ rate for alternating MDA. The main technical contribution is an infinite-dimensional dual space analysis that relates Bregman divergences on measures to dual Bregman divergences on spaces of bounded continuous functions, allowing us to control asymmetric commutator terms created by alternating updates. The results substantially generalize prior analyses restricted to bilinear objectives and also apply to nonlinear convex-concave problems on measure spaces, thereby providing a unified theoretical foundation for MDA in mean-field min-max optimization.

math.OC↗

On propagation of chaos for the Fisher-Rao gradient flow in entropic mean-field optimization

We consider a class of optimization problems on the space of probability measures motivated by the mean-field approach to studying neural networks. Such problems can be solved by constructing continuous-time gradient flows that converge to the minimizer of the energy function under consideration, and then implementing discrete-time algorithms that approximate the flow. In this work, we focus on the Fisher-Rao gradient flow and we construct an interacting particle system that approximates the flow as its mean-field limit. We discuss the connection between the energy function, the gradient flow and the particle system and explain different approaches to smoothing out the energy function with an appropriate kernel in a way that allows for the particle system to be well-defined. We provide a rigorous proof of the existence and uniqueness of thus obtained kernelized flows, as well as a propagation of chaos result that provides a theoretical justification for using the corresponding kernelized particle systems as approximation algorithms in entropic mean-field optimization.

math.OC↗

Non-convex entropic mean-field optimization via Best Response flow

We study the problem of minimizing non-convex functionals on the space of probability measures, regularized by the relative entropy (KL divergence) with respect to a fixed reference measure, as well as the corresponding problem of solving entropy-regularized non-convex-non-concave min-max problems. We utilize the Best Response flow (also known in the literature as the fictitious play flow) and study how its convergence is influenced by the relation between the degree of non-convexity of the functional under consideration, the regularization parameter and the tail behaviour of the reference measure. In particular, we demonstrate how to choose the regularizer, given the non-convex functional, so that the Best Response operator becomes a contraction with respect to the $L^1$-Wasserstein distance, which ensures the existence of its unique fixed point that is then shown to be the unique global minimizer for our optimization problem. This extends recent results where the Best Response flow was applied to solve convex optimization problems regularized by the relative entropy with respect to arbitrary reference measures, and with arbitrary values of the regularization parameter. Our results explain precisely how the assumption of convexity can be relaxed, at the expense of making a specific choice of the regularizer. Additionally, we demonstrate how these results can be applied in reinforcement learning in the context of policy optimization for Markov Decision Processes and Markov games with softmax parametrized policies in the mean-field regime.

math.OC↗

Entropic mean-field min-max problems via Best Response flow

We investigate the convergence properties of a continuous-time optimization method, the \textit{Mean-Field Best Response} flow, for solving convex-concave min-max games with entropy regularization. We introduce suitable Lyapunov functions to establish exponential convergence to the unique mixed Nash equilibrium. Additionally, we demonstrate the convergence of the fictitious play flow as a by-product of our analysis.

math.OC↗

Geometric ergodicity of modified Euler schemes for SDEs with super-linearity

As a well-known fact, the classical Euler scheme works merely for SDEs with coefficients of linear growth. In this paper, we study a general framework of modified Euler schemes, which is applicable to SDEs with super-linear drifts and encompasses numerical methods such as the tamed Euler scheme and the truncated Euler scheme. On the one hand, by exploiting an approach based on the refined basic coupling, we show that all Euler recursions within our proposed framework are geometrically ergodic under a mixed probability distance (i.e., the total variation distance plus the $L^1$-Wasserstein distance) and the weighted total variation distance. On the other hand, by utilizing the coupling by reflection, we demonstrate that the tamed Euler scheme is geometrically ergodic under the $L^1$-Wasserstein distance. In addition, as an important application, we provide a quantitative $L^1$-Wasserstein error bound between the exact invariant probability measure of an SDE with super-linearity, and the invariant probability measure of the tamed Euler scheme which is its numerical counterpart.

math.PR↗

A Fisher-Rao gradient flow for entropic mean-field min-max games

Gradient flows play a substantial role in addressing many machine learning problems. We examine the convergence in continuous-time of a \textit{Fisher-Rao} (Mean-Field Birth-Death) gradient flow in the context of solving convex-concave min-max games with entropy regularization. We propose appropriate Lyapunov functions to demonstrate convergence with explicit rates to the unique mixed Nash equilibrium.

math.OC↗

$L^2$-Wasserstein contraction for Euler schemes of elliptic diffusions and interacting particle systems

We show the $L^2$-Wasserstein contraction for the transition kernel of a discretised diffusion process, under a contractivity at infinity condition on the drift and a sufficiently high diffusivity requirement. This extends recent results that, under similar assumptions on the drift but without the diffusivity restrictions, showed the $L^1$-Wasserstein contraction, or $L^p$-Wasserstein bounds for $p > 1$ that were, however, not true contractions. We explain how showing the true $L^2$-Wasserstein contraction is crucial for obtaining the local Poincaré inequality for the transition kernel of the Euler scheme of a diffusion. Moreover, we discuss other consequences of our contraction results, such as concentration inequalities and convergence rates in KL-divergence and total variation. We also study the corresponding $L^2$-Wasserstein contraction for discretisations of interacting diffusions. As a particular application, this allows us to analyse the behaviour of particle systems that can be used to approximate a class of McKean-Vlasov SDEs that were recently studied in the mean-field optimization literature.

math.PR↗

Polyak-Łojasiewicz inequality on the space of measures and convergence of mean-field birth-death processes

The Polyak-Lojasiewicz inequality (PLI) in $\mathbb{R}^d$ is a natural condition for proving convergence of gradient descent algorithms. In the present paper, we study an analogue of PLI on the space of probability measures $\mathcal{P}(\mathbb{R}^d)$ and show that it is a natural condition for showing exponential convergence of a class of birth-death processes related to certain mean-field optimization problems. We verify PLI for a broad class of such problems for energy functions regularised by the KL-divergence.

math.OC↗

Optimal Markovian coupling for finite activity Lévy processes

We study optimal Markovian couplings of Markov processes, where the optimality is understood in terms of minimization of concave transport costs between the time-marginal distributions of the coupled processes. We provide explicit constructions of such optimal couplings for one-dimensional finite-activity Lévy processes (continuous-time random walks) whose jump distributions are unimodal but not necessarily symmetric. Remarkably, the optimal Markovian coupling does not depend on the specific concave transport cost. To this end, we combine McCann's results on optimal transport and Rogers' results on random walks with a novel uniformization construction that allows us to characterize all Markovian couplings of finite-activity Lévy processes. In particular, we show that the optimal Markovian coupling for finite-activity Lévy processes with non-symmetric unimodal Lévy measures has to allow for non-simultaneous jumps of the two coupled processes.

math.PR↗

Multi-index Antithetic Stochastic Gradient Algorithm

Stochastic Gradient Algorithms (SGAs) are ubiquitous in computational statistics, machine learning and optimisation. Recent years have brought an influx of interest in SGAs, and the non-asymptotic analysis of their bias is by now well-developed. However, relatively little is known about the optimal choice of the random approximation (e.g mini-batching) of the gradient in SGAs as this relies on the analysis of the variance and is problem specific. While there have been numerous attempts to reduce the variance of SGAs, these typically exploit a particular structure of the sampled distribution by requiring a priori knowledge of its density's mode. It is thus unclear how to adapt such algorithms to non-log-concave settings. In this paper, we construct a Multi-index Antithetic Stochastic Gradient Algorithm (MASGA) whose implementation is independent of the structure of the target measure and which achieves performance on par with Monte Carlo estimators that have access to unbiased samples from the distribution of interest. In other words, MASGA is an optimal estimator from the mean square error-computational cost perspective within the class of Monte Carlo estimators. We prove this fact rigorously for log-concave settings and verify it numerically for some examples where the log-concavity assumption is not satisfied.

stat.ML↗

Strict Kantorovich contractions for Markov chains and Euler schemes with general noise

We study contractions of Markov chains on general metric spaces with respect to some carefully designed distance-like functions, which are comparable to the total variation and the standard $L^p$-Wasserstein distances for $p \ge 1$. We present explicit lower bounds of the corresponding contraction rates. By employing the refined basic coupling and the coupling by reflection, the results are applied to Markov chains whose transitions include additive stochastic noises that are not necessarily isotropic. This can be useful in the study of Euler schemes for SDEs driven by Lévy noises. In particular, motivated by recent works on the use of heavy tailed processes in Markov Chain Monte Carlo, we show that chains driven by the $α$-stable noise can have better contraction rates than corresponding chains driven by the Gaussian noise, due to the heavy tails of the $α$-stable distribution.

math.PR↗

Exponential ergodicity for SDEs and McKean-Vlasov processes with Lévy noise

We study stochastic differential equations (SDEs) of McKean-Vlasov type with distribution dependent drifts and driven by pure jump Lévy processes. We prove a uniform in time propagation of chaos result, providing quantitative bounds on convergence rate of interacting particle systems with Lévy noise to the corresponding McKean-Vlasov SDE. By applying techniques that combine couplings, appropriately constructed $L^1$-Wasserstein distances and Lyapunov functions, we show exponential convergence of solutions of such SDEs to their stationary distributions. Our methods allow us to obtain results that are novel even for a broad class of Lévy-driven SDEs with distribution independent coefficients.

math.PR↗

Approximation of heavy-tailed distributions via stable-driven SDEs

Constructions of numerous approximate sampling algorithms are based on the well-known fact that certain Gibbs measures are stationary distributions of ergodic stochastic differential equations (SDEs) driven by the Brownian motion. However, for some heavy-tailed distributions it can be shown that the associated SDE is not exponentially ergodic and that related sampling algorithms may perform poorly. A natural idea that has recently been explored in the machine learning literature in this context is to make use of stochastic processes with heavy tails instead of the Brownian motion. In this paper we provide a rigorous theoretical framework for studying the problem of approximating heavy-tailed distributions via ergodic SDEs driven by symmetric (rotationally invariant) $α$-stable processes.

math.PR↗

Transportation inequalities for non-globally dissipative SDEs with jumps via Malliavin calculus and coupling

By using the mirror coupling for solutions of SDEs driven by pure jump Lévy processes, we extend some transportation and concentration inequalities, which were previously known only in the case where the coefficients in the equation satisfy a global dissipativity condition. Furthermore, by using the mirror coupling for the jump part and the coupling by reflection for the Brownian part, we extend analogous results for jump diffusions. To this end, we improve some previous results concerning such couplings and show how to combine the jump and the Brownian case. As a crucial step in our proof, we develop a novel method of bounding Malliavin derivatives of solutions of SDEs with both jump and Gaussian noise, which involves the coupling technique and which might be of independent interest. The bounds we obtain are new even in the case of diffusions without jumps.

math.PR↗

Non-asymptotic bounds for sampling algorithms without log-concavity

Discrete time analogues of ergodic stochastic differential equations (SDEs) are one of the most popular and flexible tools for sampling high-dimensional probability measures. Non-asymptotic analysis in the $L^2$ Wasserstein distance of sampling algorithms based on Euler discretisations of SDEs has been recently developed by several authors for log-concave probability distributions. In this work we replace the log-concavity assumption with a log-concavity at infinity condition. We provide novel $L^2$ convergence rates for Euler schemes, expressed explicitly in terms of problem parameters. From there we derive non-asymptotic bounds on the distance between the laws induced by Euler schemes and the invariant laws of SDEs, both for schemes with standard and with randomised (inaccurate) drifts. We also obtain bounds for the hierarchy of discretisation, which enables us to deploy a multi-level Monte Carlo estimator. Our proof relies on a novel construction of a coupling for the Markov chains that can be used to control both the $L^1$ and $L^2$ Wasserstein distances simultaneously. Finally, we provide a weak convergence analysis that covers both the standard and the randomised (inaccurate) drift case. In particular, we reveal that the variance of the randomised drift does not influence the rate of weak convergence of the Euler scheme to the SDE.

math.PR↗

Multilevel Monte Carlo methods for the approximation of invariant measures of stochastic differential equations

We develop a framework that allows the use of the multi-level Monte Carlo (MLMC) methodology (Giles2015) to calculate expectations with respect to the invariant measure of an ergodic SDE. In that context, we study the (over-damped) Langevin equations with a strongly concave potential. We show that, when appropriate contracting couplings for the numerical integrators are available, one can obtain a uniform in time estimate of the MLMC variance in contrast to the majority of the results in the MLMC literature. As a consequence, a root mean square error of $\mathcal{O}(\varepsilon)$ is achieved with $\mathcal{O}(\varepsilon^{-2})$ complexity on par with Markov Chain Monte Carlo (MCMC) methods, which however can be computationally intensive when applied to large data sets. Finally, we present a multi-level version of the recently introduced Stochastic Gradient Langevin Dynamics (SGLD) method (Welling and Teh, 2011) built for large datasets applications. We show that this is the first stochastic gradient MCMC method with complexity $\mathcal{O}(\varepsilon^{-2}|\log {\varepsilon}|^{3})$, in contrast to the complexity $\mathcal{O}(\varepsilon^{-3})$ of currently available methods. Numerical experiments confirm our theoretical findings.

math.NA↗