SearcharxivSearch

arXiv subjects

Michelle Carey

Publications and source records attributed to Michelle Carey.

6 recordsLinked to original sources

Statistical Analysis of Network Collections Using Persistent Homology and Functional Data Analysis

Statistical analysis of collections of networks, where each network is treated as the primary unit of observation, is of growing importance across a wide range of application domains, including gene regulatory, social, and financial networks. As networks consist of vertices and edges that do not naturally reside in Euclidean space, the direct application of conventional statistical methodologies, such as the computation of means and covariances, principal component analysis, and hypothesis testing, to samples of networks is not straightforward. A central challenge lies in defining meaningful measures of similarity or distance between networks of potentially varying sizes and structural types (e.g., directed, undirected, weighted or unweighted), particularly when no predefined node correspondence exists. To address these challenges, we introduce a framework termed functional topological data analysis (funTDA), which integrates tools from functional data analysis and topological data analysis to facilitate exploratory data analysis and inference on samples of networks. The proposed framework enables the computation of summary statistics, including means and variances, and supports the application of principal component analysis and hypothesis testing to topological features extracted from network data. Through simulation studies involving networks with varying connectivity structures, we demonstrate the ability of funTDA to distinguish between distinct network configurations. The methodology is illustrated through two real-data applications: networks constructed from pairwise word co-occurrences in novels by Jane Austen and Charles Dickens, and gene regulatory networks derived from gene expression measurements for seventeen individuals exposed to H3N2 influenza. In both applications, differences in network topology are assessed using principal component analysis and hypothesis testing.

stat.ME

Estimation of Multivariate Functional Principal Components from Sparse Functional Data

Traditional Functional Principal Component Analysis typically focuses on densely observed univariate functional data, yet many applications, particularly in longitudinal studies, involve multivariate functional data observed sparsely and irregularly across subjects. A common approach for extracting multivariate functional principal components in such settings relies on an eigen decomposition of univariate functional principal component scores to capture cross-component correlations. We propose a new approach for the estimation of multivariate functional principal components by improving the univariate eigenanalysis through maximum likelihood estimation combined with a modified Gram-Schmidt orthonormalization. The performance of the proposed approach is evaluated against two established methods, and its practical utility is demonstrated through an application to longitudinal cognitive biomarker data from an Alzheimer's disease study and a collection of data on dairy milk yield and milk compositions from research dairy farms in Ireland.

stat.ME

Estimation of Functional Principal Components from Sparse Functional Data

Sparse functional data arise when measurements are observed infrequently and at irregular time points for each subject, often in the presence of measurement error. These characteristics introduce additional challenges for functional principal component analysis. In this paper, we propose a new approach for extracting functional principal components from such data by combining basis expansion with maximum likelihood estimation. Orthogonality of the estimated eigenfunctions is preserved throughout the optimization using modified Gram-Schmidt orthonormalization. An information criterion is proposed to select both the optimal number of basis functions and the rank of the covariance structure. Principal component scores are subsequently estimated via conditional expectation, enabling accurate reconstruction of the underlying functional trajectories across the full domain despite sparse observations. Simulation studies demonstrate the effectiveness of the proposed method and show that it performs favorably compared with existing approaches. Its practical utility is illustrated through applications to CD4 cell count data from the Multicenter AIDS Cohort Study and somatic cell count data from Irish research dairy cattle. Supplementary materials, including technical details, additional simulation results, and the R package mGSFPCA, are available online.

stat.ME

Spatial Transcriptomics Iterative Hierarchical Clustering (stIHC): A Novel Method for Identifying Spatial Gene Co-Expression Modules

Recent advancements in spatial transcriptomics technologies allow researchers to simultaneously measure RNA expression levels for hundreds to thousands of genes while preserving spatial information within tissues, providing critical insights into spatial gene expression patterns, tissue organization, and gene functionality. However, existing methods for clustering spatially variable genes (SVGs) into co-expression modules often fail to detect rare or unique spatial expression patterns. To address this, we present spatial transcriptomics iterative hierarchical clustering (stIHC), a novel method for clustering SVGs into co-expression modules, representing groups of genes with shared spatial expression patterns. Through three simulations and applications to spatial transcriptomics datasets from technologies such as 10x Visium, 10x Xenium, and Spatial Transcriptomics, stIHC outperforms clustering approaches used by popular SVG detection methods, including SPARK, SPARK-X, MERINGUE, and SpatialDE. Gene Ontology enrichment analysis confirms that genes within each module share consistent biological functions, supporting the functional relevance of spatial co-expression. Robust across technologies with varying gene numbers and spatial resolution, stIHC provides a powerful tool for decoding the spatial organization of gene expression and the functional structure of complex tissues.

stat.ME

The Functional Gait Deviation Index

A typical gait analysis requires the examination of the motion of nine joint angles on the left-hand side and six joint angles on the right-hand side across multiple subjects. Due to the quantity and complexity of the data, it is useful to calculate the amount by which a subject's gait deviates from an average normal profile and to represent this deviation as a single number. Such a measure can quantify the overall severity of a condition affecting walking, monitor progress, or evaluate the outcome of an intervention prescribed to improve the gait pattern. The gait deviation index, gait profile score, and the overall abnormality measure are standard benchmarks for quantifying gait abnormality. However, these indices do not account for the intrinsic smoothness of the gait movement at each joint/plane and the potential co-variation between the joints/planes. Utilizing a multivariate functional principal components analysis we propose the functional gait deviation index (FGDI). FGDI accounts for the intrinsic smoothness of the gait movement at each joint/plane and the potential co-variation between the joints. We show that FGDI scales with overall gait function, provides a consistent measure of gait abnormality, and is implemented easily using an interactive web app.

stat.AP

Fast Stable Parameter Estimation for Linear Dynamical Systems

Dynamical systems describe the changes in processes that arise naturally from their underlying physical principles, such as the laws of motion or the conservation of mass, energy or momentum. These models facilitate a causal explanation for the drivers and impediments of the processes. But do they describe the behaviour of the observed data? And how can we quantify the models' parameters that cannot be measured directly? This paper addresses these two questions by providing a methodology for estimating the solution; and the parameters of linear dynamical systems from incomplete and noisy observations of the processes. The proposed procedure builds on the parameter cascading approach, where a linear combination of basis functions approximates the implicitly defined solution of the dynamical system. The systems' parameters are then estimated so that this approximating solution adheres to the data. By taking advantage of the linearity of the system, we have simplified the parameter cascading estimation procedure, and by developing a new iterative scheme, we achieve fast and stable computation. We illustrate our approach by obtaining a linear differential equation that represents real data from biomechanics. Comparing our approach with popular methods for estimating the parameters of linear dynamical systems, namely, the non-linear least-squares approach, simulated annealing, parameter cascading and smooth functional tempering reveals a considerable reduction in computation and an improved bias and sampling variance.

stat.ME