SearcharxivSearch

arXiv subjects

Jure Vogrinc

Publications and source records attributed to Jure Vogrinc.

11 recordsLinked to original sources

Metropolis Adjusted Langevin Trajectories: a robust alternative to Hamiltonian Monte Carlo

We introduce MALT: a new Metropolis adjusted sampler built upon the (kinetic) Langevin diffusion. Compared to Generalized Hamiltonian Monte Carlo (GHMC), the Metropolis correction is applied to whole Langevin trajectories, which prevents momentum flips, and allows for larger step-sizes. We argue that MALT yields a neater extension of HMC, preserving many desirable properties. We extend optimal scaling results of HMC to MALT for isotropic targets, and obtain the same scaling with respect to the dimension without additional assumptions. We show that MALT improves both the robustness to tuning and the sampling performance of HMC on anisotropic targets. We compare our approach with Randomized HMC, recently praised for its robustness. We show that, in continuous time, the Langevin diffusion achieves the fastest mixing rate for strongly log-concave targets. We then assess the accuracies of MALT, GHMC, HMC and RHMC when performing numerical integration on anisotropic targets, both on toy models and real data experiments on a Bayesian logistic regression. We show that MALT outperforms GHMC, standard HMC, and is competitive with RHMC.

stat.CO

Adaptive Tuning for Metropolis Adjusted Langevin Trajectories

Hamiltonian Monte Carlo (HMC) is a widely used sampler for continuous probability distributions. In many cases, the underlying Hamiltonian dynamics exhibit a phenomenon of resonance which decreases the efficiency of the algorithm and makes it very sensitive to hyperparameter values. This issue can be tackled efficiently, either via the use of trajectory length randomization (RHMC) or via partial momentum refreshment. The second approach is connected to the kinetic Langevin diffusion, and has been mostly investigated through the use of Generalized HMC (GHMC). However, GHMC induces momentum flips upon rejections causing the sampler to backtrack and waste computational resources. In this work we focus on a recent algorithm bypassing this issue, named Metropolis Adjusted Langevin Trajectories (MALT). We build upon recent strategies for tuning the hyperparameters of RHMC which target a bound on the Effective Sample Size (ESS) and adapt it to MALT, thereby enabling the first user-friendly deployment of this algorithm. We construct a method to optimize a sharper bound on the ESS and reduce the estimator variance. Easily compatible with parallel implementation, the resultant Adaptive MALT algorithm is competitive in terms of ESS rate and hits useful tradeoffs in memory usage when compared to GHMC, RHMC and NUTS.

stat.CO

Optimal design of the Barker proposal and other locally-balanced Metropolis-Hastings algorithms

We study the class of first-order locally-balanced Metropolis--Hastings algorithms introduced in Livingstone & Zanella (2021). To choose a specific algorithm within the class the user must select a balancing function $g:\mathbb{R} \to \mathbb{R}$ satisfying $g(t) = tg(1/t)$, and a noise distribution for the proposal increment. Popular choices within the class are the Metropolis-adjusted Langevin algorithm and the recently introduced Barker proposal. We first establish a universal limiting optimal acceptance rate of 57% and scaling of $n^{-1/3}$ as the dimension $n$ tends to infinity among all members of the class under mild smoothness assumptions on $g$ and when the target distribution for the algorithm is of the product form. In particular we obtain an explicit expression for the asymptotic efficiency of an arbitrary algorithm in the class, as measured by expected squared jumping distance. We then consider how to optimise this expression under various constraints. We derive an optimal choice of noise distribution for the Barker proposal, optimal choice of balancing function under a Gaussian noise distribution, and optimal choice of first-order locally-balanced algorithm among the entire class, which turns out to depend on the specific target distribution. Numerical simulations confirm our theoretical findings and in particular show that a bi-modal choice of noise distribution in the Barker proposal gives rise to a practical algorithm that is consistently more efficient than the original Gaussian version.

stat.CO

A Metropolis-class sampler for targets with non-convex support

We aim to improve upon the exploration of the general-purpose random walk Metropolis algorithm when the target has non-convex support $A \subset \mathbb{R}^d$, by reusing proposals in $A^c$ which would otherwise be rejected. The algorithm is Metropolis-class and under standard conditions the chain satisfies a strong law of large numbers and central limit theorem. Theoretical and numerical evidence of improved performance relative to random walk Metropolis are provided. Issues of implementation are discussed and numerical examples, including applications to global optimisation and rare event sampling, are presented.

math.PR

Hopping between distant basins

We present the Basin Hopping with Skipping (BH-S) algorithm for stochastic optimisation, which replaces the perturbation step of basin hopping (BH) with a so-called skipping proposal from the rare-event sampling literature. Empirical results on benchmark optimisation surfaces demonstrate that BH-S can improve performance relative to BH by encouraging non-local exploration, that is, by hopping between distant basins.

math.OC

Counterexamples for optimal scaling of Metropolis-Hastings chains with rough target densities

For sufficiently smooth targets of product form it is known that the variance of a single coordinate of the proposal in RWM (Random walk Metropolis) and MALA (Metropolis adjusted Langevin algorithm) should optimally scale as $n^{-1}$ and as $n^{-\frac{1}{3}}$ with dimension $n$, and that the acceptance rates should be tuned to $0.234$ and $0.574$. We establish counterexamples to demonstrate that smoothness assumptions of the order of $\mathcal{C}^1(\mathbb{R})$ for RWM and $\mathcal{C}^3(\mathbb{R})$ for MALA are indeed required if these scaling rates are to hold. The counterexamples identify classes of marginal targets for which these guidelines are violated, obtained by perturbing a standard Normal density (at the level of the potential for RWM and the second derivative of the potential for MALA) using roughness generated by a path of fractional Brownian motion with Hurst exponent $H$. For such targets there is strong evidence that RWM and MALA proposal variances should optimally be scaled as $n^{-\frac{1}{H}}$ and as $n^{-\frac{1}{2+H}}$ and will then obey anomalous acceptance rate guidelines. Useful heuristics resulting from this theory are discussed. The paper develops a framework capable of tackling optimal scaling results for quite general Metropolis-Hastings algorithms (possibly depending on a random environment).

math.PR

Asymptotic variance for Random walk Metropolis chains in high dimensions: logarithmic growth via the Poisson equation

There are two ways of speeding up MCMC algorithms: (1) construct more complex samplers that use gradient and higher order information about the target and (2) design a control variate to reduce the asymptotic variance. While the efficiency of (1) as a function of dimension has been studied extensively, this paper provides first rigorous results linking the growth of the asymptotic variance in (2) with dimension. Specifically, we construct a control variate for a $d$-dimensional Random walk Metropolis chain with an IID target using the solution of the Poisson equation for the scaling limit in the seminal paper "Weak convergence and optimal scaling of random walk Metropolis algorithms" of Gelman, Gilks and Roberts. We prove that the asymptotic variance of the corresponding estimator is bounded above by a multiple of $\log d/d$ over the spectral gap of the chain. The proof hinges on large deviations theory, optimal Young's inequality and Berry-Esseen type bounds. Extensions of the result to non-product targets are discussed.

math.PR

Frequency violations from random disturbances: an MCMC approach

The frequency stability of power systems is increasingly challenged by various types of disturbances. In particular, the increasing penetration of renewable energy sources is increasing the variability of power generation and at the same time reducing system inertia against disturbances. In this paper we are particularly interested in understanding how rate of change of frequency (RoCoF) violations could arise from unusually large power disturbances. We devise a novel specialization, named ghost sampling, of the Metropolis-Hastings Markov Chain Monte Carlo method that is tailored to efficiently sample rare power disturbances leading to nodal frequency violations. Generating a representative random sample addresses important statistical questions such as "which generator is most likely to be disconnected due to a RoCoF violation?" or "what is the probability of having simultaneous RoCoF violations, given that a violation occurs?" Our method can perform conditional sampling from any joint distribution of power disturbances including, for instance, correlated and non-Gaussian disturbances, features which have both been recently shown to be significant in security analyses.

eess.SY

On the Poisson equation for Metropolis-Hastings chains

This paper defines an approximation scheme for a solution of the Poisson equation of a geometrically ergodic Metropolis-Hastings chain $Φ$. The approximations give rise to a natural sequence of control variates for the ergodic average $S_k(F)=(1/k)\sum_{i=1}^{k} F(Φ_i)$, where $F$ is the force function in the Poisson equation. The main result of the paper shows that the sequence of the asymptotic variances (in the CLTs for the control-variate estimators) converges to zero and gives a rate of this convergence. Numerical examples in the case of a double-well potential are discussed.

math.PR

Binomial Symbols and Prime Moduli

We try to improve a problem asked in an Indian Math Olympiad. We give a brief overview of the work done in Saikia and Vogrinc (2011) and Saikia and Vogrinc where the authors have found a periodic sequence and the length of its period, all inspired from an Olympiad problem. The main goal of the paper is to improve the main result in Saikia and Vogrinc (2011).

math.NT

A Simple Number Theoretic Result

We derive an interesting congruence relation motivated by an Indian Olympiad problem. We give three different proofs of the theorem and mention a few interesting related results.

math.NT