SearcharxivSearch

arXiv subjects

Benjamin Eltzner

Publications and source records attributed to Benjamin Eltzner.

At least 19 recordsLinked to original sources

Two-Sample Tests for Optimal Lifts, Manifold Stability and Reverse Labeling Reflection Shap

We consider a quotient of a complete Riemannian manifold modulo an isometrically and properly acting Lie group and lifts of the quotient to the manifolds in optimal position to a reference point on the manifold. With respect to the pushed forward Riemannian volume onto the quotient we derive continuity and uniqueness a.e. and smoothness to large extents also with respect to the reference point. In consequence we derive a general manifold stability theorem: the Fr\'echet mean lies in the highest dimensional stratum assumed with positive probability, and a strong law for optimal lifts. This allows to define new two-sample tests utilizing individual optimal lifts which outperform existing two-sample tests on simulated data. They also outperform existing tests on a newly derived reverse labeling reflection shape space, that is used to model filament data of microtubules within cells in a biological application.

math.ST

Constrained Shape Analysis with Applications to RNA Structure

In many applications of shape analysis, lengths between some landmarks are constrained. For instance, biomolecules often have some bond lengths and some bond angles constrained, and variation occurs only along unconstrained bonds and constrained bonds' torsions where the latter are conveniently modelled by dihedral angles. Our work has been motivated by low resolution biomolecular chain RNA where only some prominent atomic bonds can be well identified. Here, we propose a new modelling strategy for such constrained shape analysis starting with a product of polar coordinates (polypolars), where, due to constraints, for example, some radial coordinates should be omitted, leaving products of spheres (polyspheres). We give insight into these coordinates for particular cases such as five landmarks which are motivated by a practical RNA application. We also discuss distributions for polypolar coordinates and give a specific methodology with illustration when the constrained size-and-shape variables are concentrated. There are applications of this in clustering and we give some insight into a modified version of the MINT-AGE algorithm.

stat.ME

The Long Time Limit of Diffusion Means

In statistics on manifolds, the notion of the mean of a probability distribution becomes more involved than in a linear space. Several location statistics have been proposed, which reduce to the ordinary mean in Euclidean space. A relatively new family of contenders in this field are Diffusion Means, which are a one parameter family of location statistics modeled as initial points of isotropic diffusion with the diffusion time as parameter. It is natural to consider limit cases of the diffusion time parameter and it turns out that for short times the diffusion mean set approaches the intrinsic mean set. For long diffusion times, the limit is less obvious but for spheres of arbitrary dimension the diffusion mean set has been shown to converge to the extrinsic mean set. Here, we extend this result to the real projective spaces in their unique smooth isometric embedding into a linear space. We conjecture that the long time limit is always given by the extrinsic mean in the isometric embedding for connected compact symmetric spaces with unique isometric embedding.

math.ST

A Lower Bound for Estimating Fr\'echet Means

Fr\'echet means, conceptually appealing, generalize the Euclidean expectation to general metric spaces. We explore how well Fr\'echet means can be estimated from independent and identically distributed samples and uncover a fundamental limitation: In the vicinity of a probability distribution $P$ with nonunique means, independent of sample size, it is not possible to uniformly estimate Fr\'echet means below a precision determined by the diameter of the set of Fr\'echet means of $P$. Implications were previously identified for empirical plug-in estimators as part of the phenomenon \emph{finite sample smeariness}. Our findings thus confirm inevitable statistical challenges in the estimation of Fr\'echet means on metric spaces for which there exist distributions with nonunique means. Illustrating the relevance of our lower bound, examples of extrinsic, intrinsic, Procrustes, diffusion and Wasserstein means showcase either deteriorating constants or slow convergence rates of empirical Fr\'echet means for samples near the regime of nonunique means.

math.ST

Drift Models on Complex Projective Space for Electron-Nuclear Double Resonance

ENDOR spectroscopy is an important tool to determine the complicated three-dimensional structure of biomolecules and in particular enables measurements of intramolecular distances. Usually, spectra are determined by averaging the data matrix, which does not take into account the significant thermal drifts that occur in the measurement process. In contrast, we present an asymptotic analysis for the homoscedastic drift model, a pioneering parametric model that achieves striking model fits in practice and allows both hypothesis testing and confidence intervals for spectra. The ENDOR spectrum and an orthogonal component are modeled as an element of complex projective space, and formulated in the framework of generalized Fr\'echet means. To this end, two general formulations of strong consistency for set-valued Fr\'echet means are extended and subsequently applied to the homoscedastic drift model to prove strong consistency. Building on this, central limit theorems for the ENDOR spectrum are shown. Furthermore, we extend applicability by taking into account a phase noise contribution leading to the heteroscedastic drift model. Both drift models offer improved signal-to-noise ratio over pre-existing models.

math.ST

Diffusion Means in Geometric Spaces

We introduce a location statistic for distributions on non-linear geometric spaces, the diffusion mean, serving as an extension and an alternative to the Fr\'echet mean. The diffusion mean arises as the generalization of Gaussian maximum likelihood analysis to non-linear spaces by maximizing the likelihood of a Brownian motion. The diffusion mean depends on a time parameter $t$, which admits the interpretation of the allowed variance of the diffusion. The diffusion $t$-mean of a distribution $X$ is the most likely origin of a Brownian motion at time $t$, given the end-point distribution $X$. We give a detailed description of the asymptotic behavior of the diffusion estimator and provide sufficient conditions for the diffusion estimator to be strongly consistent. Particularly, we present a smeary central limit theorem for diffusion means and we show that joint estimation of the mean and diffusion variance rules out smeariness in all directions simultaneously in general situations. Furthermore, we investigate properties of the diffusion mean for distributions on the sphere $\mathbb S^n$. Experimentally, we consider simulated data and data from magnetic pole reversals, all indicating similar or improved convergence rate compared to the Fr\'echet mean. Here, we additionally estimate $t$ and consider its effects on smeariness and uniqueness of the diffusion mean for distributions on the sphere.

math.ST

Analyzing cross-talk between superimposed signals: Vector norm dependent hidden Markov models and applications to ion channels

We propose and investigate a hidden Markov model (HMM) for the analysis of dependent, aggregated, superimposed two-state signal recordings. A major motivation for this work is that often these signals cannot be observed individually but only their superposition. Among others, such models are in high demand for the understanding of cross-talk between ion channels, where each single channel cannot be measured separately. As an essential building block, we introduce a parameterized vector norm dependent Markov chain model and characterize it in terms of permutation invariance as well as conditional independence. This building block leads to a hidden Markov chain sum process which can be used for analyzing the dependence structure of superimposed two-state signal observations within an HMM. Notably, the model parameters of the vector norm dependent Markov chain are uniquely determined by the parameters of the sum process and are therefore identifiable. We provide algorithms to estimate the parameters, discuss model selection and apply our methodology to real-world ion channel data from the heart muscle, where we show competitive gating.

stat.ME

Smeariness Begets Finite Sample Smeariness

Fréchet means are indispensable for nonparametric statistics on non-Euclidean spaces. For suitable random variables, in some sense, they "sense" topological and geometric structure. In particular, smeariness seems to indicate the presence of positive curvature. While smeariness may be considered more as an academical curiosity, occurring rarely, it has been recently demonstrated that finite sample smeariness (FSS) occurs regularly on circles, tori and spheres and affects a large class of typical probability distributions. FSS can be well described by the modulation measuring the quotient of rescaled expected sample mean variance and population variance. Under FSS it is larger than one - that is its value on Euclidean spaces - and this makes quantile based tests using tangent space approximations inapplicable. We show here that near smeary probability distributions there are always FSS probability distributions and as a first step towards the conjecture that all compact spaces feature smeary distributions, we establish directional smeariness under curvature bounds.

math.ST

Finite Sample Smeariness on Spheres

Finite Sample Smeariness (FSS) has been recently discovered. It means that the distribution of sample Fréchet means of underlying rather unsuspicious random variables can behave as if it were smeary for quite large regimes of finite sample sizes. In effect classical quantile-based statistical testing procedures do not preserve nominal size, they reject too often under the null hypothesis. Suitably designed bootstrap tests, however, amend for FSS. On the circle it has been known that arbitrarily sized FSS is possible, and that all distributions with a nonvanishing density feature FSS. These results are extended to spheres of arbitrary dimension. In particular all rotationally symmetric distributions, not necessarily supported on the entire sphere feature FSS of Type I. While on the circle there is also FSS of Type II it is conjectured that this is not possible on higher-dimensional spheres.

math.ST

Diffusion Means and Heat Kernel on Manifolds

We introduce diffusion means as location statistics on manifold data spaces. A diffusion mean is defined as the starting point of an isotropic diffusion with a given diffusivity. They can therefore be defined on all spaces on which a Brownian motion can be defined and numerical calculation of sample diffusion means is possible on a variety of spaces using the heat kernel expansion. We present several classes of spaces, for which the heat kernel is known and sample diffusion means can therefore be calculated. As an example, we investigate a classic data set from directional statistics, for which the sample Fréchet mean exhibits finite sample smeariness.

stat.ME

Clustering Schemes on the Torus with Application to RNA Clashes

Molecular structures of RNA molecules reconstructed from X-ray crystallography frequently contain errors. Motivated by this problem we examine clustering on a torus since RNA shapes can be described by dihedral angles. A previously developed clustering method for torus data involves two tuning parameters and we assess clustering results for different parameter values in relation to the problem of so-called RNA clashes. This clustering problem is part of the dynamically evolving field of statistics on manifolds. Statistical problems on the torus highlight general challenges for statistics on manifolds. Therefore, the torus PCA and clustering methods we propose make an important contribution to directional statistics and statistics on manifolds in general.

q-bio.BM

M-Variance Asymptotics and Uniqueness of Descriptors

Asymptotic theory for M-estimation problems usually focuses on the asymptotic convergence of the sample descriptor, defined as the minimizer of the sample loss function. Here, we explore a related question and formulate asymptotic theory for the minimum value of sample loss, the M-variance. Since the loss function value is always a real number, the asymptotic theory for the M-variance is comparatively simple. M-variance often satisfies a standard central limit theorem, even in situations where the asymptotics of the descriptor is more complicated as for example in case of smeariness, or if no asymptotic distribution can be given as can be the case if the descriptor space is a general metric space. We use the asymptotic results for the M-variance to formulate a hypothesis test to systematically determine for a given sample whether the underlying population loss function may have multiple global minima. We discuss three applications of our test to data, each of which presents a typical scenario in which non-uniqueness of descriptors may occur. These model scenarios are the mean on a non-euclidean space, non-linear regression and Gaussian mixture clustering.

math.ST

Geometrical Smeariness -- A new Phenomenon of Fréchet Means

In the past decades, the central limit theorem (CLT) has been generalized to non-Euclidean data spaces. Some years ago, it was found that for some random variables on the circle, the sample Fréchet mean fluctuates around the population mean asymptotically at a scale $n^{-τ}$ with exponent $τ< 1/2$ with a non-normal distribution if the probability density at the antipodal point of the mean is $\frac{1}{2π}$. The author and his collaborator recently discovered that $τ= 1/6$ for some random variables on higher dimensional spheres. In this article we show that, even more surprisingly, the phenomenon on spheres of higher dimension is qualitatively different from that on the circle, as it depends purely on geometrical properties of the space, namely its curvature, and not on the density at the antipodal point. This gives rise to the new concept of geometrical smeariness. In consequence, the sphere can be deformed, say, by removing a neighborhood of the antipodal point of the mean and gluing a flat space there, with a smooth transition piece. This yields smeariness on a manifold, which is diffeomorphic to Euclidean space. We give an example family of random variables with 2-smeary mean, i.e. with $τ= 1/6$, whose range has a hole containing the cut locus of the mean. The hole size exhibits a curse of dimensionality as it can increase with dimension, converging to the whole hemisphere opposite a local Fréchet mean. We observe smeariness in simulated landmark shapes on Kendall pre-shape space and in real data of geomagnetic north pole positions on the two-dimensional sphere.

math.ST

Finite Sample Smeariness of Fr\'echet Means and Application to Climate

Fr\'echet means on non-Euclidean spaces may exhibit nonstandard asymptotic rates rendering quantile-based asymptotic inference inapplicable. We show here that this affects, among others, all circular distributions whose support exceeds a half circle. We exhaustively describe this phenomenon and introduce a new concept which we call finite samples smeariness (FSS). In the presence of FSS, it turns out that quantile-based tests for equality of Fr\'echet means systematically feature effective levels higher than their nominal level which perseveres asymptotically in case of Type I FSS. In contrast, suitable bootstrap-based tests correct for FSS and asymptotically attain the correct level. For illustration of the relevance of FSS in real data, we apply our method to directional wind data from two European cities. It turns out that quantile based tests, not correcting for FSS, find a multitude of significant wind changes. This multitude condenses to a few years featuring significant wind changes, when our bootstrap tests are applied, correcting for FSS.

stat.ME

Stability of the Cut Locus and a Central Limit Theorem for Fréchet Means of Riemannian Manifolds

We obtain a Central Limit Theorem for closed Riemannian manifolds, clarifying along the way the geometric meaning of some of the hypotheses in Bhattacharya and Lin's Omnibus Central Limit Theorem for Fréchet means. We obtain our CLT assuming certain stability hypothesis for the cut locus, which always holds when the manifold is compact but may not be satisfied in the non-compact case.

math.DG

A Smeary Central Limit Theorem for Manifolds with Application to High Dimensional Spheres

The (CLT) central limit theorems for generalized Frechet means (data descriptors assuming values in stratified spaces, such as intrinsic means, geodesics, etc.) on manifolds from the literature are only valid if a certain empirical process of Hessians of the Frechet function converges suitably, as in the proof of the prototypical BP-CLT (Bhattacharya and Patrangenaru (2005)). This is not valid in many realistic scenarios and we provide for a new very general CLT. In particular this includes scenarios where, in a suitable chart, the sample mean fluctuates asymptotically at a scale $n^α$ with exponents $α < 1/2$ with a non-normal distribution. As the BP-CLT yields only fluctuations that are, rescaled with $n^{1/2}$ , asymptotically normal, just as the classical CLT for random vectors, these lower rates, somewhat loosely called smeariness, had to date been observed only on the circle (Hotz and Huckemann (2015)). We make the concept of smeariness on manifolds precise, give an example for two-smeariness on spheres of arbitrary dimension, and show that smeariness, although "almost never" occurring, may have serious statistical implications on a continuum of sample scenarios nearby. In fact, this effect increases with dimension, striking in particular in high dimension low sample size scenarios.

math.ST

Backward Nested Descriptors Asymptotics with Inference on Stem Cell Differentiation

For sequences of random backward nested subspaces as occur, say, in dimension reduction for manifold or stratified space valued data, asymptotic results are derived. In fact, we formulate our results more generally for backward nested families of descriptors (BNFD). Under rather general conditions, asymptotic strong consistency holds. Under additional, still rather general hypotheses, among them existence of a.s. local twice differentiable charts, asymptotic joint normality of a BNFD can be shown. If charts factor suitably, this leads to individual asymptotic normality for the last element, a principal nested mean or a principal nested geodesic, say. It turns out that these results pertain to principal nested spheres (PNS) and principal nested great subsphere (PNGS) analysis by Jung et al. (2010) as well as to the intrinsic mean on a first geodesic principal component (IMo1GPC) for manifolds and Kendall's shape spaces. A nested bootstrap two-sample test is derived and illustrated with simulations. In a study on real data, PNGS is applied to track early human mesenchymal stem cell differentiation over a coarse time grid and, among others, to locate a change point with direct consequences for the design of further studies.

math.ST

Principal Sub-manifolds

We propose a novel method of finding principal components in multivariate data sets that lie on an embedded nonlinear Riemannian manifold within a higher-dimensional space. Our aim is to extend the geometric interpretation of PCA, while being able to capture non-geodesic modes of variation in the data. We introduce the concept of a principal sub-manifold, a manifold passing through a reference point, and at any point on the manifold extending in the direction of highest variation in the space spanned by the eigenvectors of the local tangent space PCA. Compared to recent work for the case where the sub-manifold is of dimension one Panaretos et al. (2014)$-$essentially a curve lying on the manifold attempting to capture one-dimensional variation$-$the current setting is much more general. The principal sub-manifold is therefore an extension of the principal flow, accommodating to capture higher dimensional variation in the data. We show the principal sub-manifold yields the ball spanned by the usual principal components in Euclidean space. By means of examples, we illustrate how to find, use and interpret a principal sub-manifold and we present an application in shape analysis.

stat.ME