SearcharxivSearch

arXiv subjects

Anya Katsevich

Publications and source records attributed to Anya Katsevich.

12 recordsLinked to original sources

High-dimensional Laplace asymptotics up to the concentration threshold

We study high-dimensional Laplace-type integrals $I(\lambda):=(\lambda/2\pi)^{d/2}\int_{\mathbb R^d} g(x)e^{-\lambda f(x)}dx$ in the regime where both $d$ and $\lambda$ are large. Existing rigorous Laplace-expansion results in growing dimension are largely confined to the "Gaussian-approximation" regime $d^2/\lambda\to0$, which excludes many practically relevant settings that lie beyond this threshold but still satisfy the concentration condition $d/\lambda\to0$. We close this gap by deriving an explicit asymptotic expansion for $\log I(\lambda)$ with quantitative remainder bounds that remain valid throughout this intermediate region, arbitrarily close to the concentration threshold. Fix $L\ge1$ and assume that, in a neighborhood of the global minimizer of $f$, the operator norms of derivatives of $f$ and $g$ are bounded independently of $d,\lambda$ up to orders $2L+2$ and $2L$, respectively. Assuming also some mild global growth conditions, we prove $$\log I(\lambda)=\sum_{k=1}^{L-1} b_k(f,g)\lambda^{-k}+O(d^{L+1}/\lambda^L), \qquad d^{L+1}/\lambda^L\to0,$$ with coefficients satisfying $b_k(f,g)=O(d^{k+1})$. Moreover, the $b_k(f,g)$ coincide with the coefficients from the formal cumulant expansion of $\log I(\lambda)$. We also study computation for concentrating densities $\pi(x)\propto e^{-\lambda f(x)}$. For smooth observables $g$, our expansion yields closed-form, analytic approximations of $\mathbb E_{X\sim\pi}[g(X)]$. For sampling, we construct explicit polynomial transports $x_L$ such that $\pi_L:=(x_L)_\# N(0,\lambda^{-1}I_d)$ satisfies $\mathrm{TV}(\pi,\pi_L)\lesssim d^{L+1}/\lambda^L$ for $L=1,2,3,\dots$, yielding an accurate procedure arbitrarily close to the concentration threshold $d=o(\lambda)$.

math.CA

Asymptotic analysis of rare events in high dimensions

Understanding rare events is critical across domains ranging from signal processing to reliability and structural safety, extreme-weather forecasting, and insurance. The analysis of rare events is a computationally challenging problem, particularly in high dimensions $d$. In this work, we develop the first asymptotic high-dimensional theory of rare events. First, we exploit asymptotic integral methods recently developed by the first author to provide an asymptotic expansion of rare event probabilities. The expansion employs the geometry of the rare event boundary and the local behavior of the log probability density. Generically, the expansion is valid if $d^2\ll\lambda$, where $\lambda$ characterizes the extremity of the event. We prove this condition is necessary by constructing an example in which the first-order remainder is bounded above and below by $d^2/\lambda$. We also provide a nonasymptotic remainder bound which specifies the precise dependence of the remainder on $d$, $\lambda$, the density, and the boundary, and which shows that in certain cases, the condition $d^2\ll \lambda$ can be relaxed. As an application of the theory, we derive asymptotic approximations to rare probabilities under the standard Gaussian density in high dimensions. In the second part of our work, we provide an asymptotic approximation to densities conditional on rare events. This gives rise to simple procedure for approximately sampling conditionally on the rare event using independent Gaussian and exponential random variables.

math.PR

A unified theory of the high-dimensional Laplace approximation with application to Bayesian inverse problems

The Laplace approximation (LA) to posteriors is a ubiquitous tool to simplify Bayesian computation, particularly in the high-dimensional settings arising in Bayesian inverse problems. Precisely quantifying the LA accuracy is a challenging problem in the high-dimensional regime. We develop a theory of the LA accuracy to high-dimensional posteriors which both subsumes and unifies a number of results in the literature. The primary advantage of our theory is that we introduce a new degree of flexibility, which can be used to obtain problem-specific upper bounds which are much tighter than previous "rigid" bounds. We demonstrate the theory in a prototypical example of a Bayesian inverse problem, in which this flexibility enables us to improve on prior bounds by an order of magnitude. Our optimized bounds in this setting are dimension-free, and therefore valid in arbitrarily high dimensions.

math.ST

The Laplace asymptotic expansion in high dimensions

We prove that the classical Laplace asymptotic expansion (AE) of $\int_{\mathbb R^d} g(x)e^{-nu(x)}dx$, $n\gg1$ extends to the high-dimensional regime in which $d$ may grow large with $n$. More specifically, we use new techniques suitable to high-$d$ to derive an AE which formally coincides with the classical one because the terms are the same, but which now has a new small parameter. Namely under classical assumptions on $z$ and $g$ and additional bounds on the growth of $\|\nabla^kz\|$ and $\|\nabla^kg\|$ with $d$, we show the new small parameter is $d^2/n$, in the sense that $|\text{Rem}_L|\leq C_L(d^2/n)^L$ for each $L=1,2,3,\dots$, where $\text{Rem}_L$ is the $L$th order remainder. As an example, we show that the derivative bounds are satisfied with high probability for a random function $z$ arising in a standard statistical model. We also show that if the derivative bounds are relaxed, then we still obtain a valid AE in powers of a "larger" small parameter. To prove these results, we derive a very general nonasymptotic bound on $\text{Rem}_L$ which is explicit in its dependence on $g,z,d,n$. The bound holds with nearly no apriori restrictions on the magnitude of the derivative norms. We show the bound is tight for each $L$ by proving a matching lower bound for a quartic $z$ and $g\equiv1$. When $d,z,g$ are fixed and $n\to\infty$, our bound shows that $\text{Rem}_L=O(n^{-L})$. Thus our work subsumes the classical theory of the Laplace expansion, and significantly extends it into the high-$d$ regime. This broadened applicability of the expansion is extremely useful for the many modern applications requiring the computation of high-$d$ Laplace integrals. In settings where the expansion is already in use, our precise and explicit error bound is valuable both for numerical estimates and theoretical analysis, especially near the boundary of applicability of the expansion.

math.CA

Tight Bounds on the Laplace Approximation Accuracy in High Dimensions

In Bayesian inference, a widespread technique to compute integrals against a high-dimensional posterior is to use a Gaussian proxy to the posterior known as the Laplace approximation. We address the question of accuracy of the approximation in terms of TV distance, in the regime in which dimension $d$ grows with sample size $n$. Multiple prior works have shown the requirement $d^3\ll n$ is sufficient for accuracy of the approximation. But in a recent breakthrough, Kasprzak et al, 2022 derived an upper bound scaling as $d/\sqrt n$. In this work, we further refine our understanding of the Laplace approximation error by decomposing the TV error into an $O(d/\sqrt n)$ leading order term, and an $O(d^2/n)$ remainder. This decomposition has far reaching implications: first, we use it to prove that the requirement $d^2\ll n$ cannot in general be improved by showing TV$\gtrsim d/\sqrt n$ for a posterior stemming from logistic regression with Gaussian design. Second, the decomposition provides tighter and more easily computable upper bounds on the TV error. Our result also opens the door to proving the BvM in the $d^2\ll n$ regime, and correcting the Laplace approximation to account for skew; this is pursued in two follow-up works.

math.ST

The Laplace approximation accuracy in high dimensions: a refined analysis and new skew adjustment

In Bayesian inference, making deductions about a parameter of interest requires one to sample from or compute an integral against a posterior distribution. A popular method to make these computations cheaper in high-dimensional settings is to replace the posterior with its Laplace approximation (LA), a Gaussian distribution. In this work, we derive a leading order decomposition of the LA error, a powerful technique to analyze the accuracy of the approximation more precisely than was possible before. It allows us to derive the first ever skew correction to the LA which provably improves its accuracy by an order of magnitude in the high-dimensional regime. Our approach also enables us to prove both tighter upper bounds on the standard LA and the first ever lower bounds in high dimensions. In particular, we prove that $d^2\ll n$ is in general necessary for accuracy of the LA, where $d$ is dimension and $n$ is sample size. Finally, we apply our theory in two example models: a Dirichlet posterior arising from a multinomial observation, and logistic regression with Gaussian design. In the latter setting, we prove high probability bounds on the accuracy of the LA and skew-corrected LA in powers of $d/\sqrt n$ alone.

math.ST

On the Approximation Accuracy of Gaussian Variational Inference

The main computational challenge in Bayesian inference is to compute integrals against a high-dimensional posterior distribution. In the past decades, variational inference (VI) has emerged as a tractable approximation to these integrals, and a viable alternative to the more established paradigm of Markov Chain Monte Carlo. However, little is known about the approximation accuracy of VI. In this work, we bound the TV error and the mean and covariance approximation error of Gaussian VI in terms of dimension and sample size. Our error analysis relies on a Hermite series expansion of the log posterior whose first terms are precisely cancelled out by the first order optimality conditions associated to the Gaussian VI optimization problem.

math.ST

Improved dimension dependence in the Bernstein von Mises Theorem via a new Laplace approximation bound

The Bernstein-von Mises theorem (BvM) gives conditions under which the posterior distribution of a parameter $\theta\in\Theta\subseteq\mathbb R^d$ based on $n$ independent samples is asymptotically normal. In the high-dimensional regime, a key question is to determine the growth rate of $d$ with $n$ required for the BvM to hold. We show that up to a model-dependent coefficient, $n\gg d^2$ suffices for the BvM to hold in two settings: arbitrary generalized linear models, which include exponential families as a special case, and multinomial data, in which the parameter of interest is an unknown probability mass functions on $d+1$ states. Our results improve on the tightest previously known condition for posterior asymptotic normality, $n\gg d^3$. Our statements of the BvM are nonasymptotic, taking the form of explicit high-probability bounds. To prove the BvM, we derive a new simple and explicit bound on the total variation distance between a measure $\pi\propto e^{-nf}$ on $\Theta\subseteq\mathbb R^d$ and its Laplace approximation.

math.ST

From local equilibrium to numerical PDE: Metropolis crystal surface dynamics in the rough scaling limit

We derive the PDE governing the hydrodynamic limit of a Metropolis rate crystal surface height process in the "rough scaling" regime introduced by Marzuola and Weare. The PDE takes the form of a continuity equation, and the expression for the current involves a numerically computed multiplicative correction term similar to a mobility. The correction accounts for the fact that, unusually, the local equilibrium distribution of the process is not a local Gibbs measure even though the global equilibrium distribution is Gibbs. We give definitive numerical evidence of this fact, originally suggested in Gao, et. al., Pure and Applied Analysis (2021). In that paper, an approximate PDE -- our PDE, but without the correction term -- was derived for the limit of the Metropolis rate process under the assumption of a local Gibbs distribution. Our main contribution is to present a numerical method to compute the corrected macroscopic current, which is given by a function of the third spatial derivative of the height profile. Our method exploits properties of the local equilibrium (LE) state of the third order finite difference process. We find that the LE state of this process is not only useful for deriving the PDE; it also enjoys nonstandard properties which are interesting in their own right. Namely, we demonstrate that the LE state is a "rough LE", a novel kind of LE state discovered in our recent work on an Arrhenius rate crystal surface process.

math.PR

The Local Equilibrium State of a Crystal Surface Jump Process in the Rough Scaling Regime

We investigate the local equilibrium (LE) distribution of a crystal surface jump process as it approaches its hydrodynamic (continuum) limit in a nonstandard scaling regime introduced by Marzuola and Weare. The atypical scaling leads to a local equilibrium state whose structure is novel, to the best of our knowledge. The distinguishing characteristic of the new, "rough" LE state is that the ensemble average of single lattice site observables do not vary smoothly across lattice sites. We investigate numerically and analytically how the rough LE state affects the convergence mechanism via three key limits, and show that by comparison, more standard, "smooth" LE states satisfy stronger versions of these limits.

math.PR

Likelihood Maximization and Moment Matching in Low SNR Gaussian Mixture Models

We derive an asymptotic expansion for the log likelihood of Gaussian mixture models (GMMs) with equal covariance matrices in the low signal-to-noise regime. The expansion reveals an intimate connection between two types of algorithms for parameter estimation: the method of moments and likelihood optimizing algorithms such as Expectation-Maximization (EM). We show that likelihood optimization in the low SNR regime reduces to a sequence of least squares optimization problems that match the moments of the estimate to the ground truth moments one by one. This connection is a stepping stone toward the analysis of EM and maximum likelihood estimation in a wide range of models. A motivating application for the study of low SNR mixture models is cryo-electron microscopy data, which can be modeled as a GMM with algebraic constraints imposed on the mixture centers. We discuss the application of our expansion to algebraically constrained GMMs, among other example models of interest.

math.ST

On De Graaf spaces of pseudoquotients

A space of pseudoquotients $\mathcal{B}(X,S)$ is defined as equivalence classes of pairs $(x,f)$, where $x$ is an element of a non-empty set $X$, $f$ is an element of $S$, a commutative semigroup of injective maps from $X$ to $X$, and $(x,f) \sim (y,g)$ if $gx=fy$. In this note we consider a generalization of this construction where the assumption of commutativity of $S$ by Ore type conditions. As in the commutative case, $X$ can be identified with a subset of $\mathcal{B}(X,S)$ and $S$ can be extended to a group $G$ of bijections on $\mathcal{B}(X,S)$. We introduce a natural topology on $\mathcal{B}(X,S)$ and show that all elements of $G$ are homeomorphisms on $\mathcal{B}(X,S)$.

math.RA