SearcharxivSearch

arXiv subjects

Adam Quinn Jaffe

Publications and source records attributed to Adam Quinn Jaffe.

13 recordsLinked to original sources

Nonparametric Riemannian Empirical Bayes, and Denoising Measurements on Manifolds

We initiate the study of nonparametric empirical Bayes denoising methods in the setting where both the latent variables and their measurements lie on a compact Riemannian manifold, and where the likelihood is a Riemannian Gaussian distribution. Our starting point is a novel Tweedie-Eddington formula for Riemannian Gaussian mixture models which identifies a certain surrogate oracle denoiser in terms of the marginal distribution of the measurements; it avoids the explicit computation of the posterior Fr\'echet mean (as required by the Bayes denoiser) via a first-order approximation, hence we refer to it as the "tangential" Bayes denoiser. We show that this surrogate oracle achieves nearly the Bayes risk in a low-noise regime, we construct a fully data-driven approximation of it using the spectral theory of the Laplace-Beltrami operator, and we establish finite-sample rates of convergence for the distance between the the surrogate oracle and its approximation. Contrasting the nearly-parametric rates from the Euclidean setting, the rates in the Riemannian setting are slower due to the singularities of the Riemannian Gaussian density at the cut locus of its Fr\'echet mean; in the special case of the circle we establish matching lower bounds which show that our proposed denoiser is minimax-optimal, and that the denoising problem exhibits a genuinely nonparametric rate of convergence. Lastly, we implement our methodology in two scientific applications: in astronomy, the sphere-valued problem of denoising the locations of gamma ray bursts; in structural biology, the torus-valued problem of denoising pairs of torsion angles of adjacent amino acids in a protein (i.e., the Ramachandran plot).

math.ST

Wasserstein-Cram\'er-Rao Theory of Unbiased Estimation

The quantity of interest in the classical Cram\'er-Rao theory of unbiased estimation (e.g., the Cram\'er-Rao lower bound, its exact attainment for exponential families, and asymptotic efficiency of maximum likelihood estimation) is the variance, which represents the instability of an estimator when its value is compared to the value for an independently-sampled data set from the same distribution. In this paper we are interested in a quantity which represents the instability of an estimator when its value is compared to the value for an infinitesimal additive perturbation of the original data set; we refer to this as the "sensitivity" of an estimator. The resulting theory of sensitivity is based on the Wasserstein geometry in the same way that the classical theory of variance is based on the Fisher-Rao (equivalently, Hellinger) geometry, and this insight allows us to determine a collection of results which are analogous to the classical case: a Wasserstein-Cram\'er-Rao lower bound for the sensitivity of any unbiased estimator, a characterization of models in which there exist unbiased estimators achieving the lower bound exactly, and some concrete results that show that the Wasserstein projection estimator achieves the lower bound asymptotically. We use these results to treat many statistical examples, sometimes revealing new optimality properties for existing estimators and other times revealing entirely new estimators.

math.ST

Coupling Theory, Optimal Transport, and Strassen's Theorem for Eventual Domination

Many results in probability (most famously, Strassen's theorem on stochastic domination), characterize some relationship between probability distributions in terms of the existence of a particular structured coupling between them. Optimal transport, and in particular Kantorovich duality, provides a common framework for these results. We use this perspective to study "eventual domination", a class of orders arising naturally in stochastic processes (including in the analysis of branching random walks, Ising models, and diffusions), which does not satisfy the topological conditions required by standard optimal transport theory. More generally, we study the connection between distributional relations and their coupling counterparts for topologically irregular preorders (e.g. equivalence relations, partial orders). To this end, we show that Strassen's theorem "nearly holds" in this topologically irregular setting but that the full theorem admits counterexamples, including for eventual domination. The core of the proof is a novel technical result in optimal transport, which shows that the Kantorovich dual problem is well-behaved for cost functions which can be written as a non-increasing limit of lower semi-continuous functions.

math.PR

Consistency and inconsistency in $k$-means clustering

A celebrated result of Pollard proves asymptotic consistency for $k$-means clustering when the population distribution has finite variance. In this work, we point out that the population-level $k$-means clustering problem is, in fact, well-posed under the weaker assumption of a finite expectation, and we investigate whether some form of asymptotic consistency holds in this setting. As we illustrate in a variety of negative results, the complete story is quite subtle; for example, the empirical $k$-means cluster centers may fail to converge even if there exists a unique set of population $k$-means cluster centers. A detailed analysis of our negative results reveals that inconsistency arises because of an extreme form of cluster imbalance, whereby the presence of outlying samples leads to some empirical $k$-means clusters possessing very few points. We then give a collection of positive results which show that some forms of asymptotic consistency, under only the assumption of finite expectation, may be recovered by imposing some a priori degree of balance among the empirical $k$-means clusters.

math.ST

Constrained Denoising, Empirical Bayes, and Optimal Transport

In latent variables models, two important goals are denoising and deconvolution: denoising aims to estimate the latent variables, whereas deconvolution aims to estimate the distribution of the latent variables. As has been recognized in the literature over the last century, these two goals are fundamentally in tension, since denoising yields a poor estimate of the distribution of the latent variables due to shrinkage, and deconvolution yields a distribution-valued estimate that carries no unit-specific information. In this paper, we provide a systematic study of denoisers, and empirical Bayes approximations thereof, which attain optimal denoising error subject to the constraint that the distribution of the denoised data matches, in some sense, the distribution of the latent variables. Our insight is that optimal transport allows practitioners to navigate the tension between denoising and deconvolution. More precisely, we propose a modular methodology that combines any suitable unconstrained empirical Bayes denoiser (arising, e.g., via $F$-modeling, $G$-modeling, conjugate-parametric models) with any suitable information about the distribution of the latent variables (e.g., its moments, support, or an approximation of the entire distribution via deconvolution) into a single denoised data set. We prove explicit rates of convergence for our proposed methodologies, and we apply the resulting methods in applications in astronomy, baseball analytics, and marketing.

stat.ME

Fr\'echet Means in Infinite Dimensions

While there exists a well-developed asymptotic theory of Fr\'echet means of random variables taking values in a general "finite-dimensional" metric space, there are only a few known results in which the random variables can take values in an "infinite-dimensional" metric space. Presently, we develop a general asymptotic theory of Fr\'echet means in some infinite-dimensional metric spaces, which allows us to recover, strengthen, and generalize most existing results; in particular, we develop novel asymptotic theory for Fr\'echet means in some infinite-dimensional metric spaces from statistical shape analysis. The core of the proof is a novel notion of weak convergence in general metric spaces for which the results can be proven via calculus of variations.

math.PR

Large Deviations Principle for Bures-Wasserstein Barycenters

We prove the large deviations principle for empirical Bures-Wasserstein barycenters of independent, identically-distributed samples of covariance matrices and covariance operators. As an application, we explore some consequences of our results for the phenomenon of dimension-free concentration of measure for Bures-Wasserstein barycenters. Our theory reveals a novel notion of exponential tilting in the Bures-Wasserstein space, which, in analogy with Cr\'amer's theorem in the Euclidean case, solves the relative entropy projection problem under a constraint on the barycenter. Notably, this method of proof is easy to adapt to other geometric settings of interest; with the same method, we obtain large deviations principles for empirical barycenters in Riemannian manifolds and the univariate Wasserstein space, and we obtain large deviations upper bounds for empirical barycenters in the general multivariate Wasserstein space. In fact, our results are the first known large deviations principles for Fr\'echet means in any non-linear metric space.

math.PR

Constructing Maximal Germ Couplings of Brownian Motions with Drift

Consider all the possible ways of coupling together two Brownian motions with the same starting position but with different drifts onto the same probability space. It is known that there exist couplings which make these processes agree for some random, positive, maximal initial length of time. Presently, we provide an explicit, elementary construction of such couplings.

math.PR

Fr\'echet Mean Set Estimation in the Hausdorff Metric, via Relaxation

This work resolves the following question in non-Euclidean statistics: Is it possible to consistently estimate the Fr\'echet mean set of an unknown population distribution, with respect to the Hausdorff metric, when given access to independent identically-distributed samples? Our affirmative answer is based on a careful analysis of the "relaxed empirical Fr\'echet mean set estimators" which identify the set of near-minimizers of the empirical Fr\'echet functional and where the amount of "relaxation" vanishes as the number of data tends to infinity. On the theoretical side, our results include exact descriptions of which relaxation rates give weak consistency and which give strong consistency, as well as a description of an estimator which (assuming only the finiteness of certain moments and a mild condition on the metric entropy of the underlying metric space) adaptively finds the fastest possible relaxation rate for strongly consistent estimation. On the applied side, we consider the problem of estimating the set of Fermat-Weber points of an unknown distribution in the space of equidistant trees endowed with the tropical projective metric; in this setting, we provide an algorithm that provably implements our adaptive estimator, and we apply this method to real phylogenetic data.

math.ST

A Strong Duality Principle for Equivalence Couplings and Total Variation

We introduce and study a notion of duality for two classes of optimization problems commonly occurring in probability theory. That is, on an abstract measurable space $(\Omega,\mathcal{F})$, we consider pairs $(E,\mathcal{G})$ where $E$ is an equivalence relation on $\Omega$ and $\mathcal{G}$ is a sub-$\sigma$-algebra of $\mathcal{G}$; we say that $(E,\mathcal{F})$ satisfies "strong duality" if $E$ is $(\mathcal{F}\otimes\mathcal{F})$-measurable and if for all probability measures $\mathbb{P},\mathbb{P}'$ on $(\Omega,\mathcal{F})$ we have $$\max_{A\in\mathcal{G}}\vert \mathbb{P}(A)-\mathbb{P}'(A)\vert = \min_{\tilde{\mathbb{P}}\in\Pi(\mathbb{P},\mathbb{P}')}(1-\tilde{\mathbb{P}}(E)),$$ where $\Pi(\mathbb{P},\mathbb{P}')$ denotes the space of couplings of $\mathbb{P}$ and $\mathbb{P}'$, and where "max" and "min" assert that the supremum and infimum are in fact achieved. The results herein give wide sufficient conditions for strong duality to hold, thereby extending a form of Kantorovich duality to a class of cost functions which are irregular from the point of view of topology but regular from the point of view of descriptive set theory. The given conditions recover or strengthen classical results, and they have novel consequences in stochastic calculus, point process theory, and random sequence simulation.

math.PR

Asymptotic Theory of Geometric and Adaptive $k$-Means Clustering

We revisit Pollard's classical result on consistency for $k$-means clustering in Euclidean space, with a focus on extensions in two directions: first, to problems where the data may come from interesting geometric settings (e.g., Riemannian manifolds, reflexive Banach spaces, or the Wasserstein space); second, to problems where some parameters are chosen adaptively from the data (e.g., $k$-medoids or elbow-method $k$-means). Towards this end, we provide a general theory which shows that all clustering procedures described above are strongly consistent. In fact, our method of proof allows us to derive many asymptotic limit theorems beyond strong consistency. We also remove all assumptions about uniqueness of the set of optimal cluster centers.

math.ST

A Zero-One Law for Virtual Markov Chains

A virtual Markov chain (VMC) is a sequence $\{X_N\}_{N=0}^{\infty}$ of Markov chains (MCs) coupled together on the same probability space such that $X_N$ has state space $\{0,1,\ldots, N\}$ and such that removing all instances of $N~+~1$ from the sample path of $X_{N+1}$ results in the sample path of $X_N$ almost surely. In this paper, we prove an exact characterization of the triviality of the $\sigma$-algebra $\bigcap_{N=0}^{\infty}\sigma(X_N,X_{N+1},\ldots)$. The main tool for doing this is a decomposition theorem that the $\sigma$-algebra generated by a VMC is equal to the $\sigma$-algebra generated by a certain countably infinite collection of independent constituent MCs. These constituents are so-called staircase MCs (SMCs), which are defined to be inhomoheneous Markov chains on the non-negative integers which transition only by holding or by jumping to a value equal to the current index. We also develop some general aspects of the theory of SMCs, including a connection with some classical but very much under-appreciated aspects of convex analysis.

math.PR

Vietoris-Rips Complexes of Regular Polygons

Persistent homology has emerged as a novel tool for data analysis in the past two decades. However, there are still very few shapes or even manifolds whose persistent homology barcodes (say of the Vietoris-Rips complex) are fully known. Towards this direction, let $P_n$ be the boundary of a regular polygon in the plane with $n$ sides; we describe the homotopy types of Vietoris-Rips complexes of $P_n$. Indeed, when $n=(k+1)!!$ is an odd double factorial, we provide a complete characterization of the homotopy types and persistent homology of the Vietoris-Rips complexes of $P_n$ up to a scale parameter $r_n$, where $r_n$ approaches the diameter of $P_n$ as $n\to\infty$. Surprisingly, these homotopy types include spheres of all dimensions. Roughly speaking, the number of higher-dimensional spheres appearing is linked to the number of equilateral (but not necessarily equiangular) stars that can be inscribed into $P_n$. As our main tool we use the recently-developed theory of cyclic graphs and winding fractions. Furthermore, we show that the Vietoris-Rips complex of an arbitrarily dense subset of $P_n$ need not be homotopy equivalent to the Vietoris-Rips complex of $P_n$ itself, and indeed, these two complexes can have different homology groups in arbitrarily high dimensions. As an application of our results, we provide a lower bound on the Gromov-Hausdorff distance between $P_n$ and the circle.

math.MG