SearcharxivSearch

arXiv subjects

Andreas Eberle

Publications and source records attributed to Andreas Eberle.

At least 19 recordsLinked to original sources

Relaxation times of non-reversible Markov processes

We develop a systematic approach to quantify $L^2$-relaxation times for non-reversible Markov processes based on the singular value gap of the generator introduced by Chatterjee. The inverse of the singular value gap is equivalent to the relaxation time of the time-averaged transition semigroup. We show that, moreover, the singular value gap of the two-point motion also provides a lower bound on the usual spectral gap of the generator, and its inverse provides upper bounds on relaxation times without time averaging for sufficiently regular initial laws. We then introduce a method for deriving lower bounds on singular value gaps for Markov processes with degenerate noise that is based on the concept of a first- and second-order collapse of the generator. It follows ideas from hypocoercivity developed in a previous series of works but is simpler and more broadly applicable. In contrast to previous results, it includes settings with non-vanishing first-order collapse, and thus applies directly to Markov chains (in continuous time), but also to diffusion processes and piecewise-deterministic Markov processes. Our approach yields sharp upper and lower bounds for several classes of examples. First applications include the proof of a conjecture by Diaconis and Miclo on a square-root speed-up for lifted random walks on abelian groups, as well as bounds on relaxation times of switching flows perturbed by noise and of non-reversible diffusion processes.

math.PR

Convergence rates of self-repellent random walks, their local time and Event Chain Monte Carlo

We study the rate of convergence to equilibrium of the self-repellent random walk and its local time process on the discrete circle $\mathbb{Z}_n$. While the self-repellent random walk alone is non-Markovian since the jump rates depend on its history via its local time, jointly considering the evolution of the local time profile and the position yields a piecewise deterministic, non-reversible Markov process. We show that this joint process can be interpreted as a second-order lift of a reversible diffusion process, the discrete stochastic heat equation with Gaussian invariant measure. In particular, we obtain a lower bound on the relaxation time of order $\Omega(n^{3/2})$. Using a flow Poincar\'e inequality, we prove an upper bound for a slightly modified dynamics of order $O(n^2)$, matching recent conjectures in the physics literature. Furthermore, since the self-repellent random walk and its local time process coincide with the Event Chain Monte Carlo algorithm for the harmonic chain, a non-reversible MCMC method, we demonstrate that the relaxation time bound confirms the recent empirical observation that Event Chain Monte Carlo algorithms can outperform traditional MCMC methods such as Hamiltonian Monte Carlo.

math.PR

Convergence of non-reversible Markov processes via lifting and flow Poincar{\'e} inequality

We propose a general approach for quantitative convergence analysis of non-reversible Markov processes, based on the concept of second-order lifts and a variational approach to hypocoercivity. To this end, we introduce the flow Poincar{\'e} inequality, a space-time Poincar{\'e} inequality along trajectories of the semigroup, and a general divergence lemma based only on the Dirichlet form of an underlying reversible diffusion. We demonstrate the versatility of our approach by applying it to a pair of run-and-tumble particles with jamming, a model from non-equilibrium statistical mechanics, and several piecewise deterministic Markov processes used in sampling applications, in particular including general stochastic jump kernels.

math.AP

Space-time divergence lemmas and optimal non-reversible lifts of diffusions on Riemannian manifolds with boundary

Non-reversible lifts reduce the relaxation time of reversible diffusions at most by a square root. For reversible diffusions on domains in Euclidean space, or, more generally, on a Riemannian manifold with boundary, non-reversible lifts are in particular given by the Hamiltonian flow on the tangent bundle, interspersed with random velocity refreshments, or perturbed by Ornstein-Uhlenbeck noise, and reflected at the boundary. In order to prove that for certain choices of parameters, these lifts achieve the optimal square-root reduction up to a constant factor, precise upper bounds on relaxation times are required. A key tool for deriving such bounds by space-time Poincar\'e inequalities is a quantitative space-time divergence lemma. Extending previous work of Cao, Lu and Wang, we establish such a divergence lemma with explicit constants for general locally convex domains with smooth boundary in Riemannian manifolds satisfying a lower, not necessarily positive, curvature bound. As a consequence, we prove optimality of the lifts described above up to a constant factor, provided the deterministic transport part of the dynamics and the noise are adequately balanced. Our results show for example that an integrated Ornstein-Uhlenbeck process on a locally convex domain with diameter $d$ achieves a relaxation time of the order $d$, whereas, in general, the Poincar\'e constant of the domain is of the order $d^2$.

math.PR

Ballistic Convergence in Hit-and-Run Monte Carlo and a Coordinate-free Randomized Kaczmarz Algorithm

Hit-and-Run is a coordinate-free Gibbs sampler, yet the quantitative advantages of its coordinate-free property remain largely unexplored beyond empirical studies. In this paper, we prove sharp estimates for the Wasserstein contraction of Hit-and-Run in Gaussian target measures via coupling methods and conclude mixing time bounds. Our results uncover ballistic and superdiffusive convergence rates in certain settings. Furthermore, we extend these insights to a coordinate-free variant of the randomized Kaczmarz algorithm, an iterative method for linear systems, and demonstrate analogous convergence rates. These findings offer new insights into the advantages and limitations of coordinate-free methods for both sampling and optimization.

math.PR

Sticky coupling as a control variate for sensitivity analysis

We present and analyze a control variate strategy based on couplings to reduce the variance of finite difference estimators of sensitivity coefficients, called transport coefficients in the physics literature. We study the bias and variance of a sticky-coupling and a synchronous-coupling based estimator as the finite difference parameter $η$ goes to zero. For diffusions with elliptic additive noise, we show that when the drift is contractive outside a compact the bias of a sticky-coupling based estimator is bounded as $η\to 0$ and its variance behaves like $η^{-1}$, compared to the standard estimator whose bias and variance behave like $η^{-1}$ and $η^{-2}$, respectively. Under the stronger assumption that the drift is contractive everywhere, we additionally show that the bias and variance of the synchronous-coupling based estimator are both bounded as $η\to 0$. Our hypotheses include overdamped Langevin dynamics with many physically relevant non-convex potentials. We illustrate our theoretical results with numerical examples, including overdamped Langevin dynamics with a highly non-convex Lennard-Jones potential to demonstrate both failure of synchronous coupling and the effectiveness of sticky coupling in the not globally contractive setting.

math.PR

Non-reversible lifts of reversible diffusion processes and relaxation times

We propose a new concept of lifts of reversible diffusion processes and show that various well-known non-reversible Markov processes arising in applications are lifts in this sense of simple reversible diffusions. Furthermore, we introduce a concept of non-asymptotic relaxation times and show that these can at most be reduced by a square root through lifting, generalising a related result in discrete time. Finally, we demonstrate how the recently developed approach to quantitative hypocoercivity based on space-time Poincar\'e inequalities can be rephrased and simplified in the language of lifts and how it can be applied to find optimal lifts.

math.PR

Discrete sticky couplings of functional autoregressive processes

In this paper, we provide bounds in Wasserstein and total variation distances between the distributions of the successive iterates of two functional autoregressive processes with isotropic Gaussian noise of the form $Y_{k+1} = \mathrm{T}_γ(Y_k) + \sqrt{γσ^2} Z_{k+1}$ and $\tilde{Y}_{k+1} = \tilde{\mathrm{T}}_γ(\tilde{Y}_k) + \sqrt{γσ^2} \tilde{Z}_{k+1}$. More precisely, we give non-asymptotic bounds on $ρ(\mathcal{L}(Y_{k}),\mathcal{L}(\tilde{Y}_k))$, where $ρ$ is an appropriate weighted Wasserstein distance or a $V$-distance, uniformly in the parameter $γ$, and on $ρ(π_γ,\tildeπ_γ)$, where $π_γ$ and $\tildeπ_γ$ are the respective stationary measures of the two processes. The class of considered processes encompasses the Euler-Maruyama discretization of Langevin diffusions and its variants. The bounds we derive are of order $γ$ as $γ\to 0$. To obtain our results, we rely on the construction of a discrete sticky Markov chain $(W_k^{(γ)})_{k \in \mathbb{N}}$ which bounds the distance between an appropriate coupling of the two processes. We then establish stability and quantitative convergence results for this process uniformly on $γ$. In addition, we show that it converges in distribution to the continuous sticky process studied in previous work. Finally, we apply our result to Bayesian inference of ODE parameters and numerically illustrate them on two particular problems.

math.PR

Asymptotic bias of inexact Markov Chain Monte Carlo methods in high dimension

Inexact Markov Chain Monte Carlo methods rely on Markov chains that do not exactly preserve the target distribution. Examples include the unadjusted Langevin algorithm (ULA) and unadjusted Hamiltonian Monte Carlo (uHMC). This paper establishes bounds on Wasserstein distances between the invariant probability measures of inexact MCMC methods and their target distributions with a focus on understanding the precise dependence of this asymptotic bias on both dimension and discretization step size. Assuming Wasserstein bounds on the convergence to equilibrium of either the exact or the approximate dynamics, we show that for both ULA and uHMC, the asymptotic bias depends on key quantities related to the target distribution or the stationary probability measure of the scheme. As a corollary, we conclude that for models with a limited amount of interactions such as mean-field models, finite range graphical models, and perturbations thereof, the asymptotic bias has a similar dependence on the step size and the dimension as for product measures.

math.PR

Sticky nonlinear SDEs and convergence of McKean-Vlasov equations without confinement

We develop a new approach to study the long time behaviour of solutions to nonlinear stochastic differential equations in the sense of McKean, as well as propagation of chaos for the corresponding mean-field particle system approximations. Our approach is based on a sticky coupling between two solutions to the equation. We show that the distance process between the two copies is dominated by a solution to a one-dimensional nonlinear stochastic differential equation with a sticky boundary at zero. This new class of equations is then analyzed carefully. In particular, we show that the dominating equation has a phase transition. In the regime where the Dirac measure at zero is the only invariant probability measure, we prove exponential convergence to equilibrium both for the one-dimensional equation, and for the original nonlinear SDE. Similarly, propagation of chaos is shown by a componentwise sticky coupling and comparison with a system of one dimensional nonlinear SDEs with sticky boundaries at zero. The approach applies to equations without confinement potential and to interaction terms that are not of gradient type.

math.PR

Mixing Time Guarantees for Unadjusted Hamiltonian Monte Carlo

We provide quantitative upper bounds on the total variation mixing time of the Markov chain corresponding to the unadjusted Hamiltonian Monte Carlo (uHMC) algorithm. For two general classes of models and fixed time discretization step size $h$, the mixing time is shown to depend only logarithmically on the dimension. Moreover, we provide quantitative upper bounds on the total variation distance between the invariant measure of the uHMC chain and the true target measure. As a consequence, we show that an $\varepsilon$-accurate approximation of the target distribution $μ$ in total variation distance can be achieved by uHMC for a broad class of models with $O\left(d^{3/4}\varepsilon^{-1/2}\log (d/\varepsilon )\right)$ gradient evaluations, and for mean field models with weak interactions with $O\left(d^{1/2}\varepsilon^{-1/2}\log (d/\varepsilon )\right)$ gradient evaluations. The proofs are based on the construction of successful couplings for uHMC that realize the upper bounds.

math.PR

Couplings for Andersen Dynamics

Andersen dynamics is a standard method for molecular simulations, and a precursor of the Hamiltonian Monte Carlo algorithm used in MCMC inference. The stochastic process corresponding to Andersen dynamics is a PDMP (piecewise deterministic Markov process) that iterates between Hamiltonian flows and velocity randomizations of randomly selected particles. Both from the viewpoint of molecular dynamics and MCMC inference, a basic question is to understand the convergence to equilibrium of this PDMP particularly in high dimension. Here we present couplings to obtain sharp convergence bounds in the Wasserstein sense that do not require global convexity of the underlying potential energy.

math.PR

Two-scale coupling for preconditioned Hamiltonian Monte Carlo in infinite dimensions

We derive non-asymptotic quantitative bounds for convergence to equilibrium of the exact preconditioned Hamiltonian Monte Carlo algorithm (pHMC) on a Hilbert space. As a consequence, explicit and dimension-free bounds for pHMC applied to high-dimensional distributions arising in transition path sampling and path integral molecular dynamics are given. Global convexity of the underlying potential energies is not required. Our results are based on a two-scale coupling which is contractive in a carefully designed distance.

math.PR

Coupling and Convergence for Hamiltonian Monte Carlo

Based on a new coupling approach, we prove that the transition step of the Hamiltonian Monte Carlo algorithm is contractive w.r.t. a carefully designed Kantorovich (L1 Wasserstein) distance. The lower bound for the contraction rate is explicit. Global convexity of the potential is not required, and thus multimodal target distributions are included. Explicit quantitative bounds for the number of steps required to approximate the stationary distribution up to a given error are a direct consequence of contractivity. These bounds show that HMC can overcome diffusive behaviour if the duration of the Hamiltonian dynamics is adjusted appropriately.

math.PR

Quantitative contraction rates for Markov chains on general state spaces

We investigate the problem of quantifying contraction coefficients of Markov transition kernels in Kantorovich ($L^1$ Wasserstein) distances. For diffusion processes, relatively precise quantitative bounds on contraction rates have recently been derived by combining appropriate couplings with carefully designed Kantorovich distances. In this paper, we partially carry over this approach from diffusions to Markov chains. We derive quantitative lower bounds on contraction rates for Markov chains on general state spaces that are powerful if the dynamics is dominated by small local moves. For Markov chains on $\mathbb{R^d}$ with isotropic transition kernels, the general bounds can be used efficiently together with a coupling that combines maximal and reflection coupling. The results are applied to Euler discretizations of stochastic differential equations with non-globally contractive drifts, and to the Metropolis adjusted Langevin algorithm for sampling from a class of probability measures on high dimensional state spaces that are not globally log-concave.

math.PR

Couplings and quantitative contraction rates for Langevin dynamics

We introduce a new probabilistic approach to quantify convergence to equilibrium for (kinetic) Langevin processes. In contrast to previous analytic approaches that focus on the associated kinetic Fokker-Planck equation, our approach is based on a specific combination of reflection and synchronous coupling of two solutions of the Langevin equation. It yields contractions in a particular Wasserstein distance, and it provides rather precise bounds for convergence to equilibrium at the borderline between the overdamped and the underdamped regime. In particular, we are able to recover kinetic behavior in terms of explicit lower bounds for the contraction rate. For example, for a rescaled double-well potential with local minima at distance $a$, we obtain a lower bound for the contraction rate of order $Ω(a^{-1})$ provided the friction coefficient is of order $Θ(a^{-1})$.

math.PR

An Elementary Approach To Uniform In Time Propagation Of Chaos

Based on a coupling approach, we prove uniform in time propagation of chaos for weakly interacting mean-field particle systems with possibly non-convex confinement and interaction potentials. The approach is based on a combination of reflection and synchronous couplings applied to the individual particles. It provides explicit quantitative bounds that significantly extend previous results for the convex case.

math.PR

A Pose-Sensitive Embedding for Person Re-Identification with Expanded Cross Neighborhood Re-Ranking

Person re identification is a challenging retrieval task that requires matching a person's acquired image across non overlapping camera views. In this paper we propose an effective approach that incorporates both the fine and coarse pose information of the person to learn a discriminative embedding. In contrast to the recent direction of explicitly modeling body parts or correcting for misalignment based on these, we show that a rather straightforward inclusion of acquired camera view and/or the detected joint locations into a convolutional neural network helps to learn a very effective representation. To increase retrieval performance, re-ranking techniques based on computed distances have recently gained much attention. We propose a new unsupervised and automatic re-ranking framework that achieves state-of-the-art re-ranking performance. We show that in contrast to the current state-of-the-art re-ranking methods our approach does not require to compute new rank lists for each image pair (e.g., based on reciprocal neighbors) and performs well by using simple direct rank list based comparison or even by just using the already computed euclidean distances between the images. We show that both our learned representation and our re-ranking method achieve state-of-the-art performance on a number of challenging surveillance image and video datasets. The code is available online at: https://github.com/pse-ecn/pose-sensitive-embedding

cs.CV