SearcharxivSearch

arXiv subjects

Nawaf Bou-Rabee

Publications and source records attributed to Nawaf Bou-Rabee.

At least 19 recordsLinked to original sources

Provable Non-Acceleration of Standard Strang Splittings of Kinetic Langevin Dynamics

The OBABO and BAOAB schemes and the other standard Strang splittings of kinetic (underdamped) Langevin dynamics are widely used Markov chain Monte Carlo algorithms. Under a suitable friction scaling, the underlying diffusion relaxes on a ballistic time scale, suggesting that these discretizations, suitably tuned, sample targets with condition number $\kappa$ in $O(\sqrt{\kappa})$ iterations. We prove that no fixed choice of step size and friction, based only on the curvature bounds and the dimension, achieves this acceleration: total variation mixing time lower bounds for OBABO show that ballistic cold-start mixing fails uniformly over the smooth strongly convex class, and the lower bounds extend, with the same orders, to BAOAB and the other four Strang splittings. The proof transfers non-acceleration from optimization to sampling. Eliminating velocity gives an exact noisy heavy-ball recursion, and by the non-acceleration theorem of Goujaud, Taylor and Dieuleveut, for every tuning either some Gaussian target has a mode with relaxation time at least of order $\kappa$, or an attracting cycle exists on a smooth potential; dilating such a potential as $U_R(x)=R^2U(x/R)$ preserves its curvature bounds and produces metastability for a number of steps exponential in $R^2$, from an initial state at Wasserstein distance $O(R)$ from equilibrium. Using contraction estimates of Leimkuhler, Paulin and Whalley and a Wasserstein-to-total-variation regularization estimate, we prove a complementary upper bound of $O(\kappa)$ steps, up to logarithmic factors, for a fixed-parameter OBABO tuning. Hence, among fixed-parameter OBABO tunings, the optimal condition-number dependence of cold-start total variation mixing over this class is linear, up to logarithmic factors. A direct Gaussian calculation also rules out fixed-parameter acceleration for the left-endpoint exponential integrator.

math.PR

On couplings for kinetic Langevin diffusions

For the kinetic Langevin diffusion and its splitting discretizations, the hypoelliptic noise structure makes the relationship between couplings and total variation (TV) bounds more subtle than in the elliptic case. We establish that, for the kinetic Langevin equation with quadratic potential, no Markovian coupling (continuous or discrete) captures the asymptotic decay rate of the TV distance between two solutions with different initial values; the canonical iterated one-shot (or sticky) coupling, for which we derive an exact contraction formula, saturates this lower bound. On the constructive side, we show that the recent sharp TV bounds obtained by Chak and Monmarch\'e admit a natural interpretation through an explicit non-Markovian coupling, built from an optimal coalescence trajectory characterized by a classical minimum-energy control problem. For the OBABO splitting scheme, this approach additionally eliminates the Hessian-Lipschitz, step-size, and final-time assumptions in the work of Chak and Monmarch\'e.

math.PR

Neural Quantum States in Mixed Precision

Scientific computing has long relied on double precision (64-bit floating point) arithmetic to guarantee accuracy in simulations of real-world phenomena. However, the growing availability of hardware accelerators such as Graphics Processing Units (GPUs) has made low-precision formats attractive due to their superior performance, reduced memory footprint, and improved energy efficiency. In this work, we investigate the role of mixed-precision arithmetic in neural-network based Variational Monte Carlo (VMC), a widely used method for solving computationally otherwise intractable quantum many-body systems. We first derive general analytical bounds on the error introduced by reduced precision on Metropolis-Hastings MCMC, and then empirically validate these bounds on the use-case of VMC. We demonstrate that significant portions of the algorithm, in particular, sampling the quantum state, can be executed in half precision without loss of accuracy. More broadly, this work provides a theoretical framework to assess the applicability of mixed-precision arithmetic in machine-learning approaches that rely on MCMC sampling. In the context of VMC, we additionally demonstrate the practical effectiveness of mixed-precision strategies, enabling more scalable and energy-efficient simulations of quantum many-body systems.

quant-ph

Tail-Sensitive KL and R\'enyi Convergence of Unadjusted Hamiltonian Monte Carlo via One-Shot Couplings

Hamiltonian Monte Carlo (HMC) algorithms are among the most widely used sampling methods in high dimensional settings, yet their convergence properties are poorly understood in divergences that quantify relative density mismatch, such as Kullback-Leibler (KL) and R\'enyi divergences. These divergences naturally govern acceptance probabilities and warm-start requirements for Metropolis-adjusted Markov chains. In this work, we develop a framework for upgrading Wasserstein convergence guarantees for unadjusted Hamiltonian Monte Carlo (uHMC) to guarantees in tail-sensitive KL and R\'enyi divergences. Our approach is based on one-shot couplings, which we use to establish a regularization property of the uHMC transition kernel. This regularization allows Wasserstein-2 mixing-time and asymptotic bias bounds to be lifted to KL divergence, and analogous Orlicz-Wasserstein bounds to be lifted to R\'enyi divergence, paralleling earlier work of Bou-Rabee and Eberle (2023) that upgrade Wasserstein-1 bounds to total variation distance via kernel smoothing. As a consequence, our results provide quantitative control of relative density mismatch, clarify the role of discretization bias in strong divergences, and yield principled guarantees relevant both for unadjusted sampling and for generating warm starts for Metropolis-adjusted Markov chains.

stat.ML

From Continuous to Discrete: a No-U-Turn Sampler for Permutations

We introduce a discrete-space analogue of the No-U-Turn sampler on the symmetric group $S_n$, yielding a locally adaptive and reversible Markov chain Monte Carlo method for $\mathrm{Mallows}(d,\sigma_0)$. Here $d:S_n\times S_n\to[0,\infty)$ is any fixed distance on $S_n$, $\sigma_0\in S_n$ is a fixed reference permutation, and the target distribution on $S_n$ has mass function $\pi(\sigma)\propto e^{-\beta d(\sigma,\sigma_0)}$ where $\beta>0$ is the inverse temperature. The construction replaces Hamiltonian trajectories with measure-preserving group-orbit exploration. A randomized dyadic expansion is used to explore a one-dimensional orbit until a probabilistic \emph{no-underrun} criterion is met, after which the next state is sampled from the explored orbit with probability proportional to the target weights. On the theory side, embedding this transition within the Gibbs self-tuning (GIST) framework provides a concise proof of reversibility. Moreover, we construct a \emph{shift coupling} for orbit segments and prove an explicit edge-wise contraction in the Cayley distance under a mild Lipschitz condition on the energy $E(\sigma)=d(\sigma,\sigma_0)$. A path-coupling argument then yields an $O(n^2\log n)$ total-variation mixing-time bound.

math.PR

Decoupling for Markov Chains

Consider a Markov chain $(X_i)_{i\ge0}$ with invariant measure $\mu$ that admits the representation $X_{i+1}=\Phi(X_i,U_i)$, where $(U_i)_{i\ge0}$ are i.i.d. random variables and $\Phi$ is a measurable map. We introduce a tangent-decoupled process $(\widetilde X_i)_{i\ge0}$ obtained by replacing $(U_i)$ with an independent copy. Conditional on the realized backbone $(X_i)$, the sequence $(f(\widetilde X_i))$ is independent. Although $(\widetilde X_i)$ is not Markovian, under the same ergodicity assumptions that ensure a law of large numbers for $(X_i)$, the empirical averages $n^{-1}\sum_{i=1}^n f(\widetilde X_i)$ converge almost surely to $\mu(f)$. In addition, for every $f\in L^2(\mu)$ and every $N\ge1$, $$ \operatorname{Var}\!\Bigl(\sum_{i=1}^N f(X_i)\Bigr) \;\le\; 2\,\operatorname{Var}\!\Bigl(\sum_{i=1}^N f(\widetilde X_i)\Bigr), $$ and therefore $\sigma_f^2 \le 2\,\widetilde\sigma_f^{\,2}$ for the corresponding time-average variance constants. The inequality requires neither reversibility nor mixing assumptions. Its proof identifies the two sequences as tangent in the sense of decoupling theory and applies the sharp $L^2$ tangent decoupling inequality of de la Pe\~na, Yao, and Alemayehu (2025).

math.PR

The Within-Orbit Adaptive Leapfrog No-U-Turn Sampler

Locally adapting parameters within Markov chain Monte Carlo methods while preserving reversibility is notoriously difficult. The success of the No-U-Turn Sampler (NUTS) largely stems from its clever local adaptation of the integration time in Hamiltonian Monte Carlo via a geometric U-turn condition. However, posterior distributions frequently exhibit multi-scale geometries with extreme variations in scale, making it necessary to also adapt the leapfrog integrator's step size locally and dynamically. Despite its practical importance, this problem has remained largely open since the introduction of NUTS by Hoffman and Gelman (2014). To address this issue, we introduce the Within-orbit Adaptive Leapfrog No-U-Turn Sampler (WALNUTS), a generalization of NUTS that adapts the leapfrog step size at fixed intervals of simulated time as the orbit evolves. At each interval, the algorithm selects the largest step size from a dyadic schedule that keeps the energy error below a user-specified threshold. Like NUTS, WALNUTS employs biased progressive state selection to favor states with positions that are further from the initial point along the orbit. Empirical evaluations on multiscale target distributions, including Neal's funnel and the Stock-Watson stochastic volatility time-series model, demonstrate that WALNUTS achieves substantial improvements in sampling efficiency and robustness compared to standard NUTS.

stat.CO

The No-Underrun Sampler: A Locally-Adaptive, Gradient-Free MCMC Method

In this work, we introduce the No-Underrun Sampler (NURS), a locally-adaptive, gradient-free Markov chain Monte Carlo method that blends ideas from Hit-and-Run and the No-U-Turn Sampler. NURS dynamically adapts to the local scale of the target distribution without requiring gradient evaluations, making it especially suitable for applications where gradients are unavailable or costly. We establish key theoretical properties, including reversibility, formal connections to Hit-and-Run and Random Walk Metropolis, Wasserstein contraction comparable to Hit-and-Run in Gaussian targets, and bounds on the total variation distance between the transition kernels of Hit-and-Run and NURS. Empirical experiments, supported by theoretical insights, illustrate the ability of NURS to sample from Neal's funnel, a challenging multi-scale distribution from Bayesian hierarchical inference.

math.ST

Ballistic Convergence in Hit-and-Run Monte Carlo and a Coordinate-free Randomized Kaczmarz Algorithm

Hit-and-Run is a coordinate-free Gibbs sampler, yet the quantitative advantages of its coordinate-free property remain largely unexplored beyond empirical studies. In this paper, we prove sharp estimates for the Wasserstein contraction of Hit-and-Run in Gaussian target measures via coupling methods and conclude mixing time bounds. Our results uncover ballistic and superdiffusive convergence rates in certain settings. Furthermore, we extend these insights to a coordinate-free variant of the randomized Kaczmarz algorithm, an iterative method for linear systems, and demonstrate analogous convergence rates. These findings offer new insights into the advantages and limitations of coordinate-free methods for both sampling and optimization.

math.PR

Mixing of the No-U-Turn Sampler and the Geometry of Gaussian Concentration

We prove that the mixing time of the No-U-Turn Sampler (NUTS), when initialized in the concentration region of the canonical Gaussian measure, scales as $d^{1/4}$, up to logarithmic factors, where $d$ is the dimension. This scaling is expected to be sharp. This result is based on a coupling argument that leverages the geometric structure of the target distribution. Specifically, concentration of measure results in a striking uniformity in NUTS' locally adapted transitions, which holds with high probability. This uniformity is formalized by interpreting NUTS as an accept/reject Markov chain, where the mixing properties for the more uniform accept chain are analytically tractable. Additionally, our analysis uncovers a previously unnoticed issue with the path length adaptation procedure of NUTS, specifically related to looping behavior, which we address in detail.

math.PR

Incorporating Local Step-Size Adaptivity into the No-U-Turn Sampler using Gibbs Self Tuning

Adapting the step size locally in the no-U-turn sampler (NUTS) is challenging because the step-size and path-length tuning parameters are interdependent. The determination of an optimal path length requires a predefined step size, while the ideal step size must account for errors along the selected path. Ensuring reversibility further complicates this tuning problem. In this paper, we present a method for locally adapting the step size in NUTS that is an instance of the Gibbs self-tuning (GIST) framework. Our approach guarantees reversibility with an acceptance probability that depends exclusively on the conditional distribution of the step size. We validate our step-size-adaptive NUTS method on Neal's funnel density and a high-dimensional normal distribution, demonstrating its effectiveness in challenging scenarios.

stat.ME

GIST: Gibbs self-tuning for locally adaptive Hamiltonian Monte Carlo

We introduce a novel and flexible framework for constructing locally adaptive Hamiltonian Monte Carlo (HMC) samplers by Gibbs sampling the algorithm's tuning parameters conditionally based on the position and momentum at each step. For adaptively sampling path lengths, our Gibbs self-tuning (GIST) approach encompasses randomized HMC, multinomial HMC, the No-U-Turn Sampler (NUTS), and the Apogee-to-Apogee Path Sampler as special cases. We exemplify the GIST framework with a novel alternative to NUTS for locally adapting path lengths, evaluated with an exact Hamiltonian for a high-dimensional, ill-conditioned Gaussian measure and with the leapfrog integrator for a suite of diverse models.

stat.CO

Randomized Runge-Kutta-Nystr\"om Methods for Unadjusted Hamiltonian and Kinetic Langevin Monte Carlo

We introduce $5/2$- and $7/2$-order $L^2$-accurate randomized Runge-Kutta-Nystr\"{o}m methods, tailored for approximating Hamiltonian flows within non-reversible Markov chain Monte Carlo samplers, such as unadjusted Hamiltonian Monte Carlo and unadjusted kinetic Langevin Monte Carlo. We establish quantitative $5/2$-order $L^2$-accuracy upper bounds under gradient and Hessian Lipschitz assumptions on the potential energy function. The numerical experiments demonstrate the superior efficiency of the proposed unadjusted samplers on a variety of well-behaved, high-dimensional target distributions.

math.NA

Nonlinear Hamiltonian Monte Carlo & its Particle Approximation

We present a nonlinear (in the sense of McKean) generalization of Hamiltonian Monte Carlo (HMC) termed nonlinear HMC (nHMC) capable of sampling from nonlinear probability measures of mean-field type. When the underlying confinement potential is $K$-strongly convex and $L$-gradient Lipschitz, and the underlying interaction potential is gradient Lipschitz, nHMC can produce an $\varepsilon$-accurate approximation of a $d$-dimensional nonlinear probability measure in $L^1$-Wasserstein distance using $O((L/K) \log(1/\varepsilon))$ steps. Owing to a uniform-in-steps propagation of chaos phenomenon, and without further regularity assumptions, unadjusted HMC with randomized time integration for the corresponding particle approximation can achieve $\varepsilon$-accuracy in $L^1$-Wasserstein distance using $O( (L/K)^{5/3} (d/K)^{4/3} (1/\varepsilon)^{8/3} \log(1/\varepsilon) )$ gradient evaluations. These mixing/complexity upper bounds are a specific case of more general results developed in the paper for a larger class of non-logconcave, nonlinear probability measures of mean-field type.

math.PR

Mixing of Metropolis-Adjusted Markov Chains via Couplings: The High Acceptance Regime

We present a coupling framework to upper bound the total variation mixing time of various Metropolis-adjusted, gradient-based Markov kernels in the `high acceptance regime'. The approach uses a localization argument to boost local mixing of the underlying unadjusted kernel to mixing of the adjusted kernel when the acceptance rate is suitably high. As an application, mixing time guarantees are developed for a non-reversible, adjusted Markov chain based on the kinetic Langevin diffusion, where little is currently understood.

math.PR

Unadjusted Hamiltonian MCMC with Stratified Monte Carlo Time Integration

A randomized time integrator is suggested for unadjusted Hamiltonian Monte Carlo (uHMC) which involves a very minor modification to the usual Verlet time integrator, and hence, is easy to implement. For target distributions of the form $\mu(dx) \propto e^{-U(x)} dx$ where $U: \mathbb{R}^d \to \mathbb{R}_{\ge 0}$ is $K$-strongly convex but only $L$-gradient Lipschitz, and initial distributions $\nu$ with finite second moment, coupling proofs reveal that an $\varepsilon$-accurate approximation of the target distribution in $L^2$-Wasserstein distance $\boldsymbol{\mathcal{W}}^2$ can be achieved by the uHMC algorithm with randomized time integration using $O\left((d/K)^{1/3} (L/K)^{5/3} \varepsilon^{-2/3} \log( \boldsymbol{\mathcal{W}}^2(\mu, \nu) / \varepsilon)^+\right)$ gradient evaluations; whereas for such rough target densities the corresponding complexity of the uHMC algorithm with Verlet time integration is in general $O\left((d/K)^{1/2} (L/K)^2 \varepsilon^{-1} \log( \boldsymbol{\mathcal{W}}^2(\mu, \nu) / \varepsilon)^+ \right)$. Metropolis-adjustable randomized time integrators are also provided.

math.PR

Mixing Time Guarantees for Unadjusted Hamiltonian Monte Carlo

We provide quantitative upper bounds on the total variation mixing time of the Markov chain corresponding to the unadjusted Hamiltonian Monte Carlo (uHMC) algorithm. For two general classes of models and fixed time discretization step size $h$, the mixing time is shown to depend only logarithmically on the dimension. Moreover, we provide quantitative upper bounds on the total variation distance between the invariant measure of the uHMC chain and the true target measure. As a consequence, we show that an $\varepsilon$-accurate approximation of the target distribution $\mu$ in total variation distance can be achieved by uHMC for a broad class of models with $O\left(d^{3/4}\varepsilon^{-1/2}\log (d/\varepsilon )\right)$ gradient evaluations, and for mean field models with weak interactions with $O\left(d^{1/2}\varepsilon^{-1/2}\log (d/\varepsilon )\right)$ gradient evaluations. The proofs are based on the construction of successful couplings for uHMC that realize the upper bounds.

math.PR

A generalized class of strongly stable and dimension-free T-RPMD integrators

Recent work shows that strong stability and dimensionality freedom are essential for robust numerical integration of thermostatted ring-polymer molecular dynamics (T-RPMD) and path-integral molecular dynamics (PIMD), without which standard integrators exhibit non-ergodicity and other pathologies [J. Chem. Phys. 151, 124103 (2019); J. Chem. Phys. 152, 104102 (2020)]. In particular, the BCOCB scheme, obtained via Cayley modification of the standard BAOAB scheme, features a simple reparametrization of the free ring-polymer sub-step that confers strong stability and dimensionality freedom and has been shown to yield excellent numerical accuracy in condensed-phase systems with large time-steps. Here, we introduce a broader class of T-RPMD numerical integrators that exhibit strong stability and dimensionality freedom, irrespective of the Ornstein-Uhlenbeck friction schedule. In addition to considering equilibrium accuracy and time-step stability as in previous work, we evaluate the integrators on the basis of their rates of convergence to equilibrium and their efficiency at evaluating equilibrium expectation values. Within the generalized class, we find BCOCB to be superior with respect to accuracy and efficiency for various configuration-dependent observables, although other integrators within the generalized class perform better for velocity-dependent quantities. Extensive numerical evidence indicates that the stated performance guarantees hold for the strongly anharmonic case of liquid water. Both analytical and numerical results indicate that BCOCB excels over other known integrators in terms of accuracy, efficiency, and stability with respect to time-step for practical applications.

physics.chem-ph