SearcharxivSearch

arXiv subjects

J. S. Marron

Publications and source records attributed to J. S. Marron.

At least 19 recordsLinked to original sources

Principal Subsimplex Analysis

Compositional data, which are data that lie in a simplex, naturally arise in many scientific domains such as geochemistry, microbiology, and economics. In such domains, obtaining sensible lower-dimensional representations and modes of variation plays an important role. A typical approach to the problem is applying a log-ratio transformation followed by principal component analysis (PCA). However, this approach has several potential weaknesses: it can amplify variation in minor variables and obscure important variation within major variables; it is not directly applicable to datasets containing zeros, and zero imputation methods can give highly variable results; it has limited ability to capture linear patterns present on the simplex. In this paper, we propose novel methods that produce nested sequences of simplices of decreasing dimensions analogous to backwards principal component analysis. These nested sequences offer both interpretable lower dimensional representations and linear modes of variation. In addition, our methods are applicable without any modification to datasets containing zeros. We demonstrate our methods on simulated data and on relative abundances of diatom species during the late Pliocene. Supplementary materials and R implementations for this article are available online.

stat.ME

Advanced Distribution Theory for Significance in Scale Space

Smoothing methods find signals in noisy data. A challenge for Statistical inference is the choice of smoothing parameter. SiZer addressed this challenge in one-dimension by detecting significant slopes across multiple scales, but was not a completely valid testing procedure. This was addressed by the development of an advanced distribution theory that ensures fully valid inference in the 1-D setting by applying extreme value theory. A two-dimensional extension of SiZer, known as Significance in Scale Space (SSS), was developed for image data, enabling the detection of both slopes and curvatures across multiple spatial scales. However, fully valid inference for 2-D SSS has remained unavailable, largely due to the more complex dependence structure of random fields. In this paper, we use a completely different probability methodology which gives an advanced distribution theory for SSS, establishing a valid hypothesis testing procedure for both slope and curvature detection. When applied to pure noise images (no true underlying signal), the proposed method controls the Type I error, whereas the original SSS identifies spurious features across scales. When signal is present, the proposed method maintains a high level of statistical power, successfully identifying important true slopes and curvatures in real data such as gamma camera images.

math.ST

Elastic Shape Analysis of Movement Data

Osteoarthritis (OA) is a highly prevalent degenerative joint disease, and the knee is the most commonly affected joint. Biomechanical factors, particularly forces exerted during walking, are often measured in modern studies of knee joint injury and OA, and understanding the relationship among biomechanics, clinical profiles, and OA has high clinical relevance. Biomechanical forces are typically represented as curves over time, but a standard practice in biomechanics research is to summarize these curves by a small number of discrete values (or landmarks). The objective of this work is to demonstrate the added value of analyzing full movement curves over conventional discrete summaries. We developed a shape-based representation of variation in full biomechanical curve data from the Intensive Diet and Exercise for Arthritis (IDEA) study (Messier et al., 2009, 2013), and demonstrated through nested model comparisons that our approach, compared to conventional discrete summaries, yields stronger associations with OA severity and OA-related clinical traits. Notably, our work is among the first to quantitatively evaluate the added value of analyzing full movement curves over conventional discrete summaries.

stat.AP

Multi-faceted Neuroimaging Data Integration via Analysis of Subspaces

Neuroimaging studies, such as the Human Connectome Project (HCP), often collect multi-faceted and multi-block data to study the complex human brain. However, these data are often analyzed in a pairwise fashion, which can hinder our understanding of how different brain-related measures interact with each other. In this study, we comprehensively analyze the multi-block HCP data using the Data Integration via Analysis of Subspaces (DIVAS) method. We integrate structural and functional brain connectivity, substance use, cognition, and genetics in an exhaustive five-block analysis. This gives rise to the important finding that genetics is the single data modality most predictive of brain connectivity, outside of brain connectivity itself. Nearly 14\% of the variation in functional connectivity (FC) and roughly 12\% of the variation in structural connectivity (SC) is attributed to shared spaces with genetics. Moreover, investigations of shared space loadings provide interpretable associations between particular brain regions and drivers of variability, such as alcohol consumption in the substance-use data block. Novel Jackstraw hypothesis tests are developed for the DIVAS framework to establish statistically significant loadings. For example, in the (FC, SC, and Substance Use) shared space, these novel hypothesis tests highlight largely negative functional and structural connections suggesting the brain's role in physiological responses to increased substance use. Furthermore, our findings have been validated using a subset of genetically relevant siblings or twins not studied in the main analysis.

q-bio.NC

Region of Interest Detection in Melanocytic Skin Tumor Whole Slide Images -- Nevus & Melanoma

Automated region of interest detection in histopathological image analysis is a challenging and important topic with tremendous potential impact on clinical practice. The deep-learning methods used in computational pathology may help us to reduce costs and increase the speed and accuracy of cancer diagnosis. We started with the UNC Melanocytic Tumor Dataset cohort that contains 160 hematoxylin and eosin whole-slide images of primary melanomas (86) and nevi (74). We randomly assigned 80% (134) as a training set and built an in-house deep-learning method to allow for classification, at the slide level, of nevi and melanomas. The proposed method performed well on the other 20% (26) test dataset; the accuracy of the slide classification task was 92.3% and our model also performed well in terms of predicting the region of interest annotated by the pathologists, showing excellent performance of our model on melanocytic skin tumors. Even though we tested the experiments on the skin tumor dataset, our work could also be extended to other medical image detection problems to benefit the clinical evaluation and diagnosis of different tumors.

eess.IV

Data Integration Via Analysis of Subspaces (DIVAS)

Modern data collection in many data paradigms, including bioinformatics, often incorporates multiple traits derived from different data types (i.e. platforms). We call this data multi-block, multi-view, or multi-omics data. The emergent field of data integration develops and applies new methods for studying multi-block data and identifying how different data types relate and differ. One major frontier in contemporary data integration research is methodology that can identify partially-shared structure between sub-collections of data types. This work presents a new approach: Data Integration Via Analysis of Subspaces (DIVAS). DIVAS combines new insights in angular subspace perturbation theory with recent developments in matrix signal processing and convex-concave optimization into one algorithm for exploring partially-shared structure. Based on principal angles between subspaces, DIVAS provides built-in inference on the results of the analysis, and is effective even in high-dimension-low-sample-size (HDLSS) situations.

stat.ME

Measure of Strength of Evidence for Visually Observed Differences between Subpopulations

For measuring the strength of visually-observed subpopulation differences, the Population Difference Criterion is proposed to assess the statistical significance of visually observed subpopulation differences. It addresses the following challenges: in high-dimensional contexts, distributional models can be dubious; in high-signal contexts, conventional permutation tests give poor pairwise comparisons. We also make two other contributions: Based on a careful analysis we find that a balanced permutation approach is more powerful in high-signal contexts than conventional permutations. Another contribution is the quantification of uncertainty due to permutation variation via a bootstrap confidence interval. The practical usefulness of these ideas is illustrated in the comparison of subpopulations of modern cancer data.

stat.ME

Powerful Significance Testing for Unbalanced Clusters

Clustering methods are popular for revealing structure in data, particularly in the high-dimensional setting common to contemporary data science. A central statistical question is, "are the clusters really there?" One pioneering method in statistical cluster validation is SigClust, but it is severely underpowered in the important setting where the candidate clusters have unbalanced sizes, such as in rare subtypes of disease. We show why this is the case, and propose a remedy that is powerful in both the unbalanced and balanced settings, using a novel generalization of k-means clustering. We illustrate the value of our method using a high-dimensional dataset of gene expression in kidney cancer patients. A Python implementation is available at https://github.com/thomaskeefe/sigclust.

stat.ME

A Partially Functional Linear Modeling Framework for Integrating Genetic, Imaging, and Clinical Data

This paper is motivated by the joint analysis of genetic, imaging, and clinical (GIC) data collected in the Alzheimer's Disease Neuroimaging Initiative (ADNI) study. We propose a regression framework based on partially functional linear regression models to map high-dimensional GIC-related pathways for Alzheimer's Disease (AD). We develop a joint model selection and estimation procedure by embedding imaging data in the reproducing kernel Hilbert space and imposing the L0 penalty for the coefficients of genetic variables. We apply the proposed method to the ADNI dataset to identify important features from tens of thousands of genetic polymorphisms (reduced from millions using a preprocessing step) and study the effects of a certain set of informative genetic variants and the baseline hippocampus surface on thirteen future cognitive scores measuring different aspects of cognitive function. We explore the shared and different heritability patterns of these cognitive scores. Analysis results suggest that both the hippocampal and genetic data have heterogeneous effects on different scores, with the trend that the value of both hippocampi is negatively associated with the severity of cognition deficits. Polygenic effects are observed for all thirteen cognitive scores. The well-known APOE4 genotype only explains a small part of cognitive function. Shared genetic etiology exists, however, greater genetic heterogeneity exists within disease classifications after accounting for the baseline diagnosis status. These analyses are useful in further investigation of functional mechanisms for AD evolution.

stat.ME

Pairwise Nonlinear Dependence Analysis of Genomic Data

In The Cancer Genome Atlas (TCGA) data set, there are many interesting nonlinear dependencies between pairs of genes that reveal important relationships and subtypes of cancer. Such genomic data analysis requires a rapid, powerful and interpretable detection process, especially in a high-dimensional environment. We study the nonlinear patterns among the expression of pairs of genes from TCGA using a powerful tool called Binary Expansion Testing. We find many nonlinear patterns, some of which are driven by known cancer subtypes, some of which are novel.

stat.AP

Jackstraw Inference for AJIVE Data Integration

In the age of big data, data integration is a critical step especially in the understanding of how diverse data types work together and work separately. Among data integration methods, the Angle-Based Joint and Individual Variation Explained (AJIVE) approach is particularly attractive because it not only studies joint behavior but also individual behavior. Typically AJIVE scores indicate important relationships between data objects, such as clusters. An important challenge is understanding which features, i.e. variables, are associated with those relationships. This challenge is addressed by the proposal of a hypothesis test for assessing statistical significance of features. The new test is inspired by the related jackstraw method developed for Principal Component Analysis. We use a high-dimensional muti-genomic cancer data set as our strong motivation and deep illustration of the methodology.

stat.AP

Region of Interest Detection in Melanocytic Skin Tumor Whole Slide Images

Automated region of interest detection in histopathological image analysis is a challenging and important topic with tremendous potential impact on clinical practice. The deep-learning methods used in computational pathology help us to reduce costs and increase the speed and accuracy of regions of interest detection and cancer diagnosis. In this work, we propose a patch-based region of interest detection method for melanocytic skin tumor whole-slide images. We work with a dataset that contains 165 primary melanomas and nevi Hematoxylin and Eosin whole-slide images and build a deep-learning method. The proposed method performs well on a hold-out test data set including five TCGA-SKCM slides (accuracy of 93.94\% in slide classification task and intersection over union rate of 41.27\% in the region of interest detection task), showing the outstanding performance of our model on melanocytic skin tumor. Even though we test the experiments on the skin tumor dataset, our work could also be extended to other medical image detection problems, such as various tumors' classification and prediction, to help and benefit the clinical evaluation and diagnosis of different tumors.

cs.CV

Trends in Northern Hemispheric Snow Presence

This paper develops a mathematical model and statistical methods to quantify trends in presence/absence observations of snow cover (not depths) and applies these in an analysis of Northern Hemispheric observations extracted from satellite flyovers during 1967-2021. A two-state Markov chain model with periodic dynamics is introduced to analyze changes in the data in a grid by grid fashion. Trends, converted to the number of weeks of snow cover lost/gained per century, are estimated for each study grid. Uncertainty margins for these trends are developed from the model and used to assess the significance of the trend estimates. Grids with questionable data quality are identified. Among trustworthy grids, snow presence is seen to be declining in almost twice as many grids as it is advancing. While Arctic and southern latitude snow presence is found to be rapidly receding, other locations, such as Eastern Canada, are experiencing advancing snow cover.

stat.AP

Statistical Significance of Clustering using Soft Thresholding

Clustering methods have led to a number of important discoveries in bioinformatics and beyond. A major challenge in their use is determining which clusters represent important underlying structure, as opposed to spurious sampling artifacts. This challenge is especially serious, and very few methods are available, when the data are very high in dimension. Statistical Significance of Clustering (SigClust) is a recently developed cluster evaluation tool for high dimensional low sample size data. An important component of the SigClust approach is the very definition of a single cluster as a subset of data sampled from a multivariate Gaussian distribution. The implementation of SigClust requires the estimation of the eigenvalues of the covariance matrix for the null multivariate Gaussian distribution. We show that the original eigenvalue estimation can lead to a test that suffers from severe inflation of type-I error, in the important case where there are a few very large eigenvalues. This paper addresses this critical challenge using a novel likelihood based soft thresholding approach to estimate these eigenvalues, which leads to a much improved SigClust. Major improvements in SigClust performance are shown by both mathematical analysis, based on the new notion of Theoretical Cluster Index, and extensive simulation studies. Applications to some cancer genomic data further demonstrate the usefulness of these improvements.

stat.ME

Scaled torus principal component analysis

A particularly challenging context for dimensionality reduction is multivariate circular data, i.e., data supported on a torus. Such kind of data appears, e.g., in the analysis of various phenomena in ecology and astronomy, as well as in molecular structures. This paper introduces Scaled Torus Principal Component Analysis (ST-PCA), a novel approach to perform dimensionality reduction with toroidal data. ST-PCA finds a data-driven map from a torus to a sphere of the same dimension and a certain radius. The map is constructed with multidimensional scaling to minimize the discrepancy between pairwise geodesic distances in both spaces. ST-PCA then resorts to principal nested spheres to obtain a nested sequence of subspheres that best fits the data, which can afterwards be inverted back to the torus. Numerical experiments illustrate how ST-PCA can be used to achieve meaningful dimensionality reduction on low-dimensional torii, particularly with the purpose of clusters separation, while two data applications in astronomy (three-dimensional torus) and molecular biology (on a seven-dimensional torus) show that ST-PCA outperforms existing methods for the investigated datasets.

stat.ME

Non-Euclidean Analysis of Joint Variations in Multi-Object Shapes

This paper considers joint analysis of multiple functionally related structures in classification tasks. In particular, our method developed is driven by how functionally correlated brain structures vary together between autism and control groups. To do so, we devised a method based on a novel combination of (1) non-Euclidean statistics that can faithfully represent non-Euclidean data in Euclidean spaces and (2) a non-parametric integrative analysis method that can decompose multi-block Euclidean data into joint, individual, and residual structures. We find that the resulting joint structure is effective, robust, and interpretable in recognizing the underlying patterns of the joint variation of multi-block non-Euclidean data. We verified the method in classifying the structural shape data collected from cases that developed and did not develop into Autistic Spectrum Disorder (ASD).

stat.ML

Persistent topology of protein space

Protein fold classification is a classic problem in structural biology and bioinformatics. We approach this problem using persistent homology. In particular, we use alpha shape filtrations to compare a topological representation of the data with a different representation that makes use of knot-theoretic ideas. We use the statistical method of Angle-based Joint and Individual Variation Explained (AJIVE) to understand similarities and differences between these representations.

q-bio.BM

New Perspectives on Centering

Data matrix centering is an ever-present yet under-examined aspect of data analysis. Functional data analysis (FDA) often operates with a default of centering such that the vectors in one dimension have mean zero. We find that centering along the other dimension identifies a novel useful mode of variation beyond those familiar in FDA. We explore ambiguities in both matrix orientation and nomenclature. Differences between centerings and their potential interaction can be easily misunderstood. We propose a unified framework and new terminology for centering operations. We clearly demonstrate the intuition behind and consequences of each centering choice with informative graphics. We also propose a new direction energy hypothesis test as part of a series of diagnostics for determining which choice of centering is best for a data set. We explore the application of these diagnostics in several FDA settings.

stat.ME