SearcharxivSearch

arXiv subjects

Janice L. Scealy

Publications and source records attributed to Janice L. Scealy.

8 recordsLinked to original sources

Principal Subsimplex Analysis

Compositional data, which are data that lie in a simplex, naturally arise in many scientific domains such as geochemistry, microbiology, and economics. In such domains, obtaining sensible lower-dimensional representations and modes of variation plays an important role. A typical approach to the problem is applying a log-ratio transformation followed by principal component analysis (PCA). However, this approach has several potential weaknesses: it can amplify variation in minor variables and obscure important variation within major variables; it is not directly applicable to datasets containing zeros, and zero imputation methods can give highly variable results; it has limited ability to capture linear patterns present on the simplex. In this paper, we propose novel methods that produce nested sequences of simplices of decreasing dimensions analogous to backwards principal component analysis. These nested sequences offer both interpretable lower dimensional representations and linear modes of variation. In addition, our methods are applicable without any modification to datasets containing zeros. We demonstrate our methods on simulated data and on relative abundances of diatom species during the late Pliocene. Supplementary materials and R implementations for this article are available online.

stat.ME

Robust Estimation of Location in Matrix Manifolds Using the Projected Frobenius Median

We propose a robust method for location estimation in various matrix manifolds based on the projected Frobenius median, which is closely related to the spatial median. This method applies broadly to matrix manifolds, including Stiefel and Grassmann manifolds, Kendall shape spaces as well as to projective Stiefel manifolds, a type of quotient space of a Stiefel manifold. Our approach involves computation of the Frobenius median in an ambient Euclidean space followed by projection onto the relevant matrix manifold. Our estimation method is computationally attractive, has a unique solution provided the sample data are not colinear in the ambient Euclidean space, has desirable robustness features and has appropriate equivariance properties under natural groups of transformations. We establish asymptotic normality under mild conditions and derive the influence function for matrix manifolds of interest. Simulation studies on the rank-1 complex Grassmann manifold and the projective Stiefel manifold further show the applicability and robustness of our method. We also apply our method to a real-world earthquake moment tensor dataset.

stat.ME

Regression for spherical responses with linear and spherical covariates using a scaled link function

We propose a regression model in which the responses are spherical variables and the covariates include linear and/or spherical variables. A novel link function is introduced by extending the Möbius transformation on the sphere. This link function is an anisotropic mapping that enables scale control along each axis of the spherical covariates and for each linear covariate. It generalizes several well-known link functions for circular or linear covariates. Each parameter of the link function is clearly interpretable. For the error distribution, we consider a general class of elliptically symmetric distributions, which includes the Kent distribution, the elliptically symmetric angular Gaussian distribution, and the scaled von Mises-Fisher distribution. Axes of symmetry of the error distribution are determined using a method involving parallel transport. Maximum likelihood estimation is feasible via reparameterization of the proposed model. Moreover, the parameters of the link function and the shape/scale parameters of the error distribution are orthogonal in the sense of the Fisher information matrix. The proposed regression model is illustrated using two real datasets. An R software package accompanies this article.

stat.ME

A Robust Extrinsic Single-index Model for Spherical Data

Regression with a spherical response is challenging due to the absence of linear structure, making standard regression models inadequate. Existing methods, mainly parametric, lack the flexibility to capture the complex relationship induced by spherical curvature, while methods based on techniques from Riemannian geometry often suffer from computational difficulties. The non-Euclidean structure further complicates robust estimation, with very limited work addressing this issue, despite the common presence of outliers in directional data. This article introduces a new semi-parametric approach, the extrinsic single-index model (ESIM) and its robust estimation, to address these limitations. We establish large-sample properties of the proposed estimator with a wide range of loss functions and assess their robustness using the influence function and standardized influence function. Specifically, we focus on the robustness of the exponential squared loss (ESL), demonstrating comparable efficiency and superior robustness over least squares loss under high concentration. We also examine how the tuning parameter for the ESL balances efficiency and robustness, providing guidance on its optimal choice. The computational efficiency and robustness of our methods are further illustrated via simulations and applications to geochemical compositional data.

stat.ME

Generalized Score Matching

Score matching is an estimation procedure that has been developed for statistical models whose probability density function is known up to proportionality but whose normalizing constant is intractable, so that maximum likelihood is difficult or impossible to implement. To date, applications of score matching have focused more on continuous IID models. Motivated by various data modelling problems, this article proposes a unified asymptotic theory of generalized score matching developed under the independence assumption, covering both continuous and discrete response data, thereby giving a sound basis for score-matchingbased inference. Real data analyses and simulation studies provide convincing evidence of strong practical performance of the proposed methods.

stat.ME

Robust score matching for compositional data

The restricted polynomially-tilted pairwise interaction (RPPI) distribution gives a flexible model for compositional data. It is particularly well-suited to situations where some of the marginal distributions of the components of a composition are concentrated near zero, possibly with right skewness. This article develops a method of tractable robust estimation for the model by combining two ideas. The first idea is to use score matching estimation after an additive log-ratio transformation. The resulting estimator is automatically insensitive to zeros in the data compositions. The second idea is to incorporate suitable weights in the estimating equations. The resulting estimator is additionally resistant to outliers. These properties are confirmed in simulation studies where we further also demonstrate that our new outlier-robust estimator is efficient in high concentration settings, even in the case when there is no model contamination. An example is given using microbiome data. A user-friendly R package accompanies the article.

stat.ME

Generalized Score Matching for Regression

Many probabilistic models that have an intractable normalizing constant may be extended to contain covariates. Since the evaluation of the exact likelihood is difficult or even impossible for these models, score matching was proposed to avoid explicit computation of the normalizing constant. In the literature, score matching has so far only been developed for models in which the observations are independent and identically distributed (IID). However, the IID assumption does not hold in the traditional fixed design setting for regression-type models. To deal with the estimation of these covariate-dependent models, this paper presents a new score matching approach for independent but not necessarily identically distributed data under a general framework for both continuous and discrete responses, which includes a novel generalized score matching method for count response regression. We prove that our proposed score matching estimators are consistent and asymptotically normal under mild regularity conditions. The theoretical results are supported by simulation studies and a real-data example. Additionally, our simulation results indicate that, compared to approximate maximum likelihood estimation, the generalized score matching produces estimates with substantially smaller biases in an application to doctoral publication data.

math.ST

Score matching for compositional distributions

Compositional data and multivariate count data with known totals are challenging to analyse due to the non-negativity and sum-to-one constraints on the sample space. It is often the case that many of the compositional components are highly right-skewed, with large numbers of zeros. A major limitation of currently available estimators for compositional models is that they either cannot handle many zeros in the data or are not computationally feasible in moderate to high dimensions. We derive a new set of novel score matching estimators applicable to distributions on a Riemannian manifold with boundary, of which the standard simplex is a special case. The score matching method is applied to estimate the parameters in a new flexible truncation model for compositional data and we show that the estimators are scalable and available in closed form. Through extensive simulation studies, the scoring methodology is demonstrated to work well for estimating the parameters in the new truncation model and also for the Dirichlet distribution. We apply the new model and estimators to real microbiome compositional data and show that the model provides a good fit to the data.

stat.ME