SearcharxivSearch

arXiv subjects

Sam Power

Publications and source records attributed to Sam Power.

At least 19 recordsLinked to original sources

Non-Uniform Random Scans in Gibbs Sampling and CAVI

Gibbs sampling and coordinate ascent variational inference (CAVI) are two basic coordinate-wise methods for statistical computation. Recent analyses under strong log-concavity establish convergence rates for versions of these algorithms that update one uniformly selected block at each step. We extend both results to arbitrary fixed, strictly positive selection probabilities. The rates are governed by a selection-adapted convexity constant $\lambda^{\star}_{\theta}$, defined using the block-smoothness constants and the selection probabilities $\theta$. The same constant yields a contraction of relative entropy for the Gibbs sampler and a contraction of the mean-field objective gap for random-scan CAVI. The new bounds recover the uniform-scan results, and are never weaker than the naive comparison based on the smallest selection probability. They provide a principled way to adapt the scan to heterogeneous block geometry using curvature information.

stat.CO

Robustness of random-walk Metropolis for steep potentials

In Markov chain Monte Carlo sampling, light-tailed target distributions present something of a poisoned chalice: their light tails offer good confinement, and tend to imply good mixing properties for natural continuous-time dynamics, but the steepness of their tail decay means that they often fall outside of the scope of modern quantitative convergence theory. For usual gradient-based samplers, this reflects a genuine instability issue, whereby Metropolis acceptance rates can degrade badly. In this work, we study the gradient-free random-walk Metropolis sampler, and show that for a wide range of light-tailed targets, the acceptance probability remains stable for reasonable choices of proposal variance, from which effective and favourable mixing time estimates can be deduced. The analysis relies on a simple relationship between the first and second derivatives of the log-density of the target distribution.

math.PR

A Complexity Bound for the Kent-Ganeiber-Mardia Sampler for the Bingham Distribution

The Bingham distribution is a family of antipodally symmetric distributions on the unit sphere, characterised by an exponential-of-quadratic change of measure with respect to the uniform distribution. Kent, Ganeiber and Mardia proposed a rejection sampler for generating samples from Bingham distributions using proposals from an angular central Gaussian (ACG) distribution. Their empirical results suggest that the least efficient regime is the high-concentration limit, where the acceptance probability is of order $d^{-1/2}$ in dimension $d$, implying a polynomial complexity guarantee. In this note, we verify this dimension-dependent prediction, establishing the uniform guarantee $\inf\{\alpha_D:D=D^\top\in\mathbb{R}^{d\times d}\}\ge c_\star/\sqrt{d}$, where $c_\star=0.759\ldots$. A one-dimensional high-concentration limit demonstrates that the $d^{-1/2}$ rate is unimprovable and that even the constant $c_\star$ cannot be improved beyond $0.857\ldots$. The proof relies on a novel interpretation of the acceptance probability and a comparison principle for weighted sums of chi-squared random variables, which may be of independent interest.

stat.CO

Wasserstein Contraction of Coordinate Ascent Variational Inference

We study the non-asymptotic contraction in Wasserstein distance of the sequential, parallel, and random-scan coordinate ascent variational inference algorithms. This is shown to hold under a functional smoothness condition of the optimality maps and a transportation-information inequality at their fixed points. Our results are sharp and general, and as opposed to those based on global strong log-concavity assumptions, they allow for local convergence on smooth, non-smooth, and discrete manifolds, including within the context of data augmentation. We consider many applications in statistical physics and Bayesian statistics. These include pairwise Markov Random field models such as Ising and Curie-Weiss, unbalanced Bayesian Gaussian Mixture Models, high-dimensional Bayesian Probit Regression, and high-dimensional Logistic Regression with P\'olya--Gamma random variables (i.e. Jaakkola-Jordan's algorithm). In many of these models, these represent the first available convergence results of their kind.

stat.ML

The sharp one-dimensional convex sub-Gaussian comparison constant

Let $X$ be an integrable real random variable with mean zero and two-sided sub-Gaussian tail $\mathbb{P}(|X|>t)\le 2e^{-t^{2}/2}$ for all $t\ge 0$. We determine the smallest constant $c_\star$ such that $X$ is dominated in convex order by $c_\star G$, where $G$ is standard normal. Equivalently, $c_\star^2$ is the sharp one-dimensional convex sub-Gaussian comparison constant appearing in the \emph{Optimization Constants in Mathematics} repository~\cite{optimization-constants-repo}. We show that $c_\star$ is given by an explicit system of one-dimensional equations and is attained by an extremal distribution that saturates the tail constraint. Numerically, $c_\star \approx 2.30952$ (so $c_\star^2 \approx 5.33386$). We also determine the analogous sharp constant under a two-sided sub-exponential tail bound, with convex domination by a scaled Laplace law. Finally, we record two higher-dimensional consequences: a sequential tensorization principle for multivariate convex domination, and a dimension-free Gaussian comparator for the cone generated by convex ridge functions (the linear convex order).

math.PR

Some aspects of robustness in modern Markov Chain Monte Carlo

Markov Chain Monte Carlo (MCMC) is a flexible approach to approximate sampling from intractable probability distributions, with a rich theoretical foundation and comprising a wealth of exemplar algorithms. While the qualitative correctness of MCMC algorithms is often easy to ensure, their practical efficiency is contingent on the `target' distribution being reasonably well-behaved. In this work, we concern ourself with the scenario in which this good behaviour is called into question, reviewing an emerging line of work on `robust' MCMC algorithms which can perform acceptably even in the face of certain pathologies. We focus on two particular pathologies which, while simple, can already have dramatic effects on standard `local' algorithms. The first is roughness, whereby the target distribution varies so rapidly that the numerical stability of the algorithm is tenuous. The second is flatness, whereby the landscape of the target distribution is instead so barren and uninformative that one becomes lost in uninteresting parts of the state space. In each case, we formulate the pathology in concrete terms, review a range of proposed algorithmic remedies to the pathology, and outline promising directions for future research.

stat.CO

Distributional Training Data Attribution: What do Influence Functions Sample?

Randomness is an unavoidable part of training deep learning models, yet something that traditional training data attribution algorithms fail to rigorously account for. They ignore the fact that, due to stochasticity in the initialisation and batching, training on the same dataset can yield different models. In this paper, we address this shortcoming through introducing distributional training data attribution (d-TDA), the goal of which is to predict how the distribution of model outputs (over training runs) depends upon the dataset. Intriguingly, we find that influence functions (IFs), a popular data attribution tool, are 'secretly distributional': they emerge from our framework as the limit to unrolled differentiation, without requiring restrictive convexity assumptions. This provides a new perspective on the effectiveness of IFs in deep learning. We demonstrate the practical utility of d-TDA in experiments, including improving data pruning for vision transformers and identifying influential examples with diffusion models.

cs.LG

Analysis of Multiple-try Metropolis via Poincar\'e inequalities

We study the Multiple-try Metropolis algorithm using the framework of Poincar\'e inequalities. We describe the Multiple-try Metropolis as an auxiliary variable implementation of a resampling approximation to an ideal Metropolis--Hastings algorithm. Under suitable moment conditions on the importance weights, we derive explicit Poincar\'e comparison results between the Multiple-try algorithm and the ideal algorithm. We characterize the spectral gap of the latter, and finally in the Gaussian case prove explicit non-asymptotic convergence bounds for Multiple-try Metropolis by comparison.

stat.CO

A New Proof of Sub-Gaussian Norm Concentration Inequality

We present a new method for proving the norm concentration inequality of sub-Gaussian variables. Our proof is based on an averaged version of the moment generating function, termed the averaged moment generating function. Our method applies to both vector cases to bound the vector norm and matrix cases to bound the operator norm. Compared with the widely adopted $\varepsilon$-net technique-based proof of the sub-Gaussian norm concentration inequality, our method does not rely on the union bound and promises a tighter concentration bound.

math.PR

Towards practical PDMP sampling: Metropolis adjustments, locally adaptive step-sizes, and NUTS-based time lengths

Piecewise-Deterministic Markov Processes (PDMPs) hold significant promise for sampling from complex probability distributions. However, their practical implementation is hindered by the need to compute model-specific bounds. Conversely, while Hamiltonian Monte Carlo (HMC) offers a generally efficient approach to sampling, its inability to adaptively tune step sizes impedes its performance when sampling complex distributions like funnels. To address these limitations, we introduce three innovative concepts: (a) a Metropolis-adjusted approximation for PDMP simulation that eliminates the need for explicit bounds without compromising the invariant measure, (b) an adaptive step size mechanism compatible with the Metropolis correction, and (c) a No U-Turn Sampler (NUTS)-inspired scheme for dynamically selecting path lengths in PDMPs. These three ideas can be seamlessly integrated into a single, `doubly-adaptive' PDMP sampler with favourable robustness and efficiency properties.

stat.CO

Prediction-Centric Uncertainty Quantification via MMD

Deterministic mathematical models, such as those specified via differential equations, are a powerful tool to communicate scientific insight. However, such models are necessarily simplified descriptions of the real world. Generalised Bayesian methodologies have been proposed for inference with misspecified models, but these are typically associated with vanishing parameter uncertainty as more data are observed. In the context of a misspecified deterministic mathematical model, this has the undesirable consequence that posterior predictions become deterministic and certain, while being incorrect. Taking this observation as a starting point, we propose Prediction-Centric Uncertainty Quantification, where a mixture distribution based on the deterministic model confers improved uncertainty quantification in the predictive context. Computation of the mixing distribution is cast as a (regularised) gradient flow of the maximum mean discrepancy (MMD), enabling consistent numerical approximations to be obtained. Results are reported on both a toy model from population ecology and a real model of protein signalling in cell biology.

stat.ME

Weak Poincar\'e inequality comparisons for ideal and hybrid slice sampling

Using the framework of weak Poincar\'e inequalities, we provide a general comparison between Hybrid and Ideal Slice Sampling in terms of their corresponding Dirichlet forms. In particular, under suitable assumptions Hybrid Slice Sampling inherits fast convergence from Ideal Slice Sampling and conversely. We apply our results to analyse the convergence of the Independent Metropolis-Hastings, stepping-out and shrinkage, as well as Hit-and-Run within slice sampling algorithms.

stat.CO

Weak Poincar\'e Inequalities for Markov chains: theory and applications

We investigate the application of Weak Poincar\'e Inequalities (WPI) to Markov chains to study their rates of convergence and to derive complexity bounds. At a theoretical level we investigate the necessity of the existence of WPIs to ensure \mathrm{L}^{2}-convergence, in particular by establishing equivalence with the Resolvent Uniform Positivity-Improving (RUPI) condition and providing a counterexample. From a more practical perspective, we extend the celebrated Cheeger's inequalities to the subgeometric setting, and further apply these techniques to study random-walk Metropolis algorithms for heavy-tailed target distributions and to obtain lower bounds on pseudo-marginal algorithms.

math.PR

Explicit convergence bounds for Metropolis Markov chains: isoperimetry, spectral gaps and profiles

We derive the first explicit bounds for the spectral gap of a random walk Metropolis algorithm on $R^d$ for any value of the proposal variance, which when scaled appropriately recovers the correct $d^{-1}$ dependence on dimension for suitably regular invariant distributions. We also obtain explicit bounds on the ${\rm L}^2$-mixing time for a broad class of models. In obtaining these results, we refine the use of isoperimetric profile inequalities to obtain conductance profile bounds, which also enable the derivation of explicit bounds in a much broader class of models. We also obtain similar results for the preconditioned Crank--Nicolson Markov chain, obtaining dimension-independent bounds under suitable assumptions.

math.PR

Adaptive Tuning for Metropolis Adjusted Langevin Trajectories

Hamiltonian Monte Carlo (HMC) is a widely used sampler for continuous probability distributions. In many cases, the underlying Hamiltonian dynamics exhibit a phenomenon of resonance which decreases the efficiency of the algorithm and makes it very sensitive to hyperparameter values. This issue can be tackled efficiently, either via the use of trajectory length randomization (RHMC) or via partial momentum refreshment. The second approach is connected to the kinetic Langevin diffusion, and has been mostly investigated through the use of Generalized HMC (GHMC). However, GHMC induces momentum flips upon rejections causing the sampler to backtrack and waste computational resources. In this work we focus on a recent algorithm bypassing this issue, named Metropolis Adjusted Langevin Trajectories (MALT). We build upon recent strategies for tuning the hyperparameters of RHMC which target a bound on the Effective Sample Size (ESS) and adapt it to MALT, thereby enabling the first user-friendly deployment of this algorithm. We construct a method to optimize a sharper bound on the ESS and reduce the estimator variance. Easily compatible with parallel implementation, the resultant Adaptive MALT algorithm is competitive in terms of ESS rate and hits useful tradeoffs in memory usage when compared to GHMC, RHMC and NUTS.

stat.CO

Comparison of Markov chains via weak Poincar\'e inequalities with application to pseudo-marginal MCMC

We investigate the use of a certain class of functional inequalities known as weak Poincar\'e inequalities to bound convergence of Markov chains to equilibrium. We show that this enables the straightforward and transparent derivation of subgeometric convergence bounds for methods such as the Independent Metropolis--Hastings sampler and pseudo-marginal methods for intractable likelihoods, the latter being subgeometric in many practical settings. These results rely on novel quantitative comparison theorems between Markov chains. Associated proofs are simpler than those relying on drift/minorization conditions and the tools developed allow us to recover and further extend known results as particular cases. We are then able to provide new insights into the practical use of pseudo-marginal algorithms, analyse the effect of averaging in Approximate Bayesian Computation (ABC) and the use of products of independent averages, and also to study the case of lognormal weights relevant to particle marginal Metropolis--Hastings (PMMH).

stat.CO

PDMP Monte Carlo methods for piecewise-smooth densities

There has been substantial interest in developing Markov chain Monte Carlo algorithms based on piecewise-deterministic Markov processes. However existing algorithms can only be used if the target distribution of interest is differentiable everywhere. The key to adapting these algorithms so that they can sample from to densities with discontinuities is defining appropriate dynamics for the process when it hits a discontinuity. We present a simple condition for the transition of the process at a discontinuity which can be used to extend any existing sampler for smooth densities, and give specific choices for this transition which work with popular algorithms such as the Bouncy Particle Sampler, the Coordinate Sampler and the Zig-Zag Process. Our theoretical results extend and make rigorous arguments that have been presented previously, for instance constructing samplers for continuous densities restricted to a bounded domain, and we present a version of the Zig-Zag Process that can work in such a scenario. Our novel approach to deriving the invariant distribution of a piecewise-deterministic Markov process with boundaries may be of independent interest.

math.ST