SearcharxivSearch

arXiv subjects

Luo Xiao

Publications and source records attributed to Luo Xiao.

18 recordsLinked to original sources

A functional joint model with baseline functional covariates: linking sitting accumulation patterns to physical function and mortality among older women

In large-scale epidemiological studies, it is often of interest to investigate joint relationships between longitudinal and time-to-event outcomes with exposures that are trajectories or functions. Our motivation study is the Objective Physical Activity and Cardiovascular Health (OPACH) Study, which collected accelerometry-measured physical activity in 6,489 older women. One of the scientific aims is to understand sedentary behavior accumulation patterns and its association with physical function (longitudinal) and mortality (time-to-event). We propose a novel approach that first converts raw accelerometry data into daily sitting bout accumulation profiles, which are treated as functional covariates, and then develop a functional joint model for longitudinal and time-to-event outcomes that incorporates a baseline functional covariate. The longitudinal process is modeled using functional data methods and linked to the survival process through functional principal component scores. Both sub-models include interpretable linear scalar-on-function regression coefficients to capture flexible dose-response associations between sitting accumulated across varying bout durations and health outcomes. Estimation is carried out via an efficient expectation-maximization (EM) algorithm with penalized spline approximations. Simulation studies demonstrate accurate parameter estimation and reliable model selection. Application to the OPACH data reveals flexible and interpretable dose-response relationships between sitting bout durations, physical function, and mortality.

stat.AP

Multivariate Multinomial Logit Model with ANOVA Decomposition for Correlated Categorical Outcomes

In medical research, patients often have multiple interdependent outcomes, such as posttraumatic stress disorder (PTSD), depression, and pain among trauma survivors. Most existing research uses multinomial regression to analyze these interdependent outcomes separately, which ignores correlations between concurrent conditions. This omission may lead to loss of information and reduced predictive accuracy. Accounting for correlations between multiple categorical outcomes requires a high-dimensional parameter space, making model estimation challenging. In this paper, we propose a multivariate multinomial logit model that captures outcome correlations and uses the ANOVA decomposition of the parameter space to reduce the number of parameters. The ANOVA decomposition enables explicit conditional model formulations, which allow for a computationally much simpler composite likelihood for model estimation. We develop an efficient Minorization-Maximization (MM) algorithm to maximize the composite likelihood, which also incorporates variable selection via a bridge penalty. Simulation studies are conducted to evaluate our method, demonstrating its accuracy in parameter estimation and variable selection. We further illustrate our method using data from the AURORA study.

stat.ME

A novel smoothing-based goodness-of-fit test of covariance for multivariate sparse functional data

Accurately specifying covariance structures is critical for valid inference in longitudinal and functional data analysis, particularly when data are sparsely observed. In this study, we develop a global goodness-of-fit test to assess parametric covariance structures in multivariate sparse functional data. Our contribution is twofold. First, we extend the univariate goodness-of-fit test proposed by Chen et al. (2019) to better accommodate sparse data by improving error variance estimation and applying positive semi-definite smoothing to covariance estimation. These corrections ensure appropriate Type I error control under sparse designs. Second, we introduce a multivariate extension of the improved test that jointly evaluates covariance structures across multiple outcomes, employing novel test statistics based on the maximum and $\ell_2$ norms to account for inter-outcome dependencies and enhance statistical power. Through extensive simulation studies, we demonstrate that the proposed methods maintain proper Type I error rates and achieve greater power than univariate tests with multiple testing adjustments. Applications to longitudinal neuroimaging and clinical data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Parkinson's Progression Marker Initiative (PPMI) illustrate the practical utility of the proposed methods for evaluating covariance structures in sparse multivariate longitudinal data.

stat.ME

Joint Linked Component Analysis for Multiview Data

In this work, we propose the joint linked component analysis (joint\_LCA) for multiview data. Unlike classic methods which extract the shared components in a sequential manner, the objective of joint\_LCA is to identify the view-specific loading matrices and the rank of the common latent subspace simultaneously. We formulate a matrix decomposition model where a joint structure and an individual structure are present in each data view, which enables us to arrive at a clean svd representation for the cross covariance between any pair of data views. An objective function with a novel penalty term is then proposed to achieve simultaneous estimation and rank selection. In addition, a refitting procedure is employed as a remedy to reduce the shrinkage bias caused by the penalization.

stat.ML

Sparse and Integrative Principal Component Analysis for Multiview Data

We consider dimension reduction of multiview data, which are emerging in scientific studies. Formulating multiview data as multi-variate data with block structures corresponding to the different views, or views of data, we estimate top eigenvectors from multiview data that have two-fold sparsity, elementwise sparsity and blockwise sparsity. We propose a Fantope-based optimization criterion with multiple penalties to enforce the desired sparsity patterns and a denoising step is employed to handle potential presence of heteroskedastic noise across different data views. An alternating direction method of multipliers (ADMM) algorithm is used for optimization. We derive the l2 convergence of the estimated top eigenvectors and establish their sparsity and support recovery properties. Numerical studies are used to illustrate the proposed method.

stat.ME

Joint Model for Survival and Multivariate Sparse Functional Data with Application to a Study of Alzheimer's Disease

Studies of Alzheimer's disease (AD) often collect multiple longitudinal clinical outcomes, which are correlated and predictive of AD progression. It is of great scientific interest to investigate the association between the outcomes and time to AD onset. We model the multiple longitudinal outcomes as multivariate sparse functional data and propose a functional joint model linking multivariate functional data to event time data. In particular, we propose a multivariate functional mixed model (MFMM) to identify the shared progression pattern and outcome-specific progression patterns of the outcomes, which enables more interpretable modeling of associations between outcomes and AD onset. The proposed method is applied to the Alzheimer's Disease Neuroimaging Initiative study (ADNI) and the functional joint model sheds new light on inference of five longitudinal outcomes and their associations with AD onset. Simulation studies also confirm the validity of the proposed model.

stat.ME

Model Testing for Generalized Scalar-on-Function Linear Models

Scalar-on-function linear models are commonly used to regress functional predictors on a scalar response. However, functional models are more difficult to estimate and interpret than traditional linear models, and may be unnecessarily complex for a data application. Hypothesis testing can be used to guide model selection by determining if a functional predictor is necessary. Using a mixed effects representation with penalized splines and variance component tests, we propose a framework for testing functional linear models with responses from exponential family distributions. The proposed method can accommodate dense and sparse functional data, and be used to test functional predictors for no effect and form of the effect. We show via simulation study that the proposed method achieves the nominal level and has high power, and we demonstrate its utility with two data applications.

stat.ME

Fast Covariance Estimation for Multivariate Sparse Functional Data

Covariance estimation is essential yet underdeveloped for analyzing multivariate functional data. We propose a fast covariance estimation method for multivariate sparse functional data using bivariate penalized splines. The tensor-product B-spline formulation of the proposed method enables a simple spectral decomposition of the associated covariance operator and explicit expressions of the resulting eigenfunctions as linear combinations of B-spline bases, thereby dramatically facilitating subsequent principal component analysis. We derive a fast algorithm for selecting the smoothing parameters in covariance smoothing using leave-one-subject-out cross-validation. The method is evaluated with extensive numerical studies and applied to an Alzheimer's disease study with multiple longitudinal outcomes.

stat.ME

An Approximate Restricted Likelihood Ratio Test for Variance Components in Generalized Linear Mixed Models

Generalized linear mixed models (GLMMs) are used to model responses from exponential families with a combination of fixed and random effects. For variance components in GLMMs, we propose an approximate restricted likelihood ratio test that conducts testing on the working responses used in penalized quasi-likelihood estimation. This presents the hypothesis test in terms of normalized responses, allowing for application of existing testing methods for linear mixed models. Our test is flexible, computationally efficient, and outperforms several competitors. We illustrate the utility of the proposed method with an extensive simulation study and two data applications. An R package is provided.

stat.ME

Fast Covariance Estimation for Sparse Functional Data

Smoothing of noisy sample covariances is an important component in functional data analysis. We propose a novel covariance smoothing method based on penalized splines and associated software. The proposed method is a bivariate spline smoother that is designed for covariance smoothing and can be used for sparse functional or longitudinal data. We propose a fast algorithm for covariance smoothing using leave-one-subject-out cross validation. Our simulations show that the proposed method compares favorably against several commonly used methods. The method is applied to a study of child growth led by one of coauthors and to a public dataset of longitudinal CD4 counts.

stat.ME

Simple fixed-effects inference for complex functional models

We propose simple inferential approaches for the fixed effects in complex functional mixed effects models. We estimate the fixed effects under the independence of functional residuals assumption and then bootstrap independent units (e.g. subjects) to estimate the variability of and conduct inference in the form of hypothesis testing on the fixed effects parameters. Simulations show excellent coverage probability of the confidence intervals and size of tests. Methods are motivated by and applied to the Baltimore Longitudinal Study of Aging (BLSA), though they are applicable to other studies that collect correlated functional data.

stat.ME

On the sample covariance matrix estimator of reduced effective rank population matrices, with applications to fPCA

This work provides a unified analysis of the properties of the sample covariance matrix $Σ_n$ over the class of $p\times p$ population covariance matrices $Σ$ of reduced effective rank $r_e(Σ)$. This class includes scaled factor models and covariance matrices with decaying spectrum. We consider $r_e(Σ)$ as a measure of matrix complexity, and obtain sharp minimax rates on the operator and Frobenius norm of $Σ_n-Σ$, as a function of $r_e(Σ)$ and $\|Σ\|_2$, the operator norm of $Σ$. With guidelines offered by the optimal rates, we define classes of matrices of reduced effective rank over which $Σ_n$ is an accurate estimator. Within the framework of these classes, we perform a detailed finite sample theoretical analysis of the merits and limitations of the empirical scree plot procedure routinely used in PCA. We show that identifying jumps in the empirical spectrum that consistently estimate jumps in the spectrum of $Σ$ is not necessarily informative for other goals, for instance for the selection of those sample eigenvalues and eigenvectors that are consistent estimates of their population counterparts. The scree plot method can still be used for selecting consistent eigenvalues, for appropriate threshold levels. We provide a threshold construction and also give a rule for checking the consistency of the corresponding sample eigenvectors. We specialize these results and analysis to population covariance matrices with polynomially decaying spectra, and extend it to covariance operators with polynomially decaying spectra. An application to fPCA illustrates how our results can be used in functional data analysis.

math.ST

Movement Prediction Using Accelerometers in a Human Population

We introduce statistical methods for predicting the types of human activity at sub-second resolution using triaxial accelerometry data. The major innovation is that we use labeled activity data from some subjects to predict the activity labels of other subjects. To achieve this, we normalize the data across subjects by matching the standing up and lying down portions of triaxial accelerometry data. This is necessary to account for differences between the variability in the position of the device relative to gravity, which are induced by body shape and size as well as by the ambiguous definition of device placement. We also normalize the data at the device level to ensure that the magnitude of the signal at rest is similar across devices. After normalization we use overlapping movelets (segments of triaxial accelerometry time series) extracted from some of the subjects to predict the movement type of the other subjects. The problem was motivated by and is applied to a laboratory study of 20 older participants who performed different activities while wearing accelerometers at the hip. Prediction results based on other people's labeled dictionaries of activity performed almost as well as those obtained using their own labeled dictionaries. These findings indicate that prediction of activity types for data collected during natural activities of daily living may actually be possible.

stat.AP

Convex Banding of the Covariance Matrix

We introduce a new sparse estimator of the covariance matrix for high-dimensional models in which the variables have a known ordering. Our estimator, which is the solution to a convex optimization problem, is equivalently expressed as an estimator which tapers the sample covariance matrix by a Toeplitz, sparsely-banded, data-adaptive matrix. As a result of this adaptivity, the convex banding estimator enjoys theoretical optimality properties not attained by previous banding or tapered estimators. In particular, our convex banding estimator is minimax rate adaptive in Frobenius and operator norms, up to log factors, over commonly-studied classes of covariance matrices, and over more general classes. Furthermore, it correctly recovers the bandwidth when the true covariance is exactly banded. Our convex formulation admits a simple and efficient algorithm. Empirical studies demonstrate its practical effectiveness and illustrate that our exactly-banded estimator works well even when the true covariance matrix is only close to a banded matrix, confirming our theoretical results. Our method compares favorably with all existing methods, in terms of accuracy and speed. We illustrate the practical merits of the convex banding estimator by showing that it can be used to improve the performance of discriminant analysis for classifying sound recordings.

math.ST

Fast Covariance Estimation for High-dimensional Functional Data

For smoothing covariance functions, we propose two fast algorithms that scale linearly with the number of observations per function. Most available methods and software cannot smooth covariance matrices of dimension $J \times J$ with $J>500$; the recently introduced sandwich smoother is an exception, but it is not adapted to smooth covariance matrices of large dimensions such as $J \ge 10,000$. Covariance matrices of order $J=10,000$, and even $J=100,000$, are becoming increasingly common, e.g., in 2- and 3-dimensional medical imaging and high-density wearable sensor data. We introduce two new algorithms that can handle very large covariance matrices: 1) FACE: a fast implementation of the sandwich smoother and 2) SVDS: a two-step procedure that first applies singular value decomposition to the data matrix and then smoothes the eigenvectors. Compared to existing techniques, these new algorithms are at least an order of magnitude faster in high dimensions and drastically reduce memory requirements. The new algorithms provide instantaneous (few seconds) smoothing for matrices of dimension $J=10,000$ and very fast ($<$ 10 minutes) smoothing for $J=100,000$. Although SVDS is simpler than FACE, we provide ready to use, scalable R software for FACE. When incorporated into R package {\it refund}, FACE improves the speed of penalized functional regression by an order of magnitude, even for data of normal size ($J <500$). We recommend that FACE be used in practice for the analysis of noisy and high-dimensional functional data.

stat.ME

On the theoretic and practical merits of the banding estimator for large covariance matrices

This paper considers the banding estimator proposed in Bickel and Levina (2008) for estimation of large covariance matrices. We prove that the banding estimator achieves rate-optimality under the operator norm, for a class of approximately banded covariance matrices, improving the existing results in Bickel and Levina (2008). In addition, we propose a Stein's unbiased risk estimate (Sure)-type approach for selecting the bandwidth for the banding estimator. Simulations indicate that the Sure-tuned banding estimator outperforms competing estimators.

math.ST

Fast Bivariate Penalized Splines: the Sandwich Smoother

We propose a fast penalized spline method for bivariate smoothing. Univariate P-spline smoothers (Eilers and Marx, 1996) are applied simultaneously along both coordinates. The new smoother has a sandwich form which suggested the name "sandwich smoother" to a referee. The sandwich smoother has a tensor product structure that simplifies an asymptotic analysis and it can be fast computed. We derive a local central limit theorem for the sandwich smoother, with simple expressions for the asymptotic bias and variance, by showing that the sandwich smoother is asymptotically equivalent to a bivariate kernel regression estimator with a product kernel. As far as we are aware, this is the first central limit theorem for a bivariate spline estimator of any type. Our simulation study shows that the sandwich smoother is orders of magnitude faster to compute than other bivariate spline smoothers, even when the latter are computed using a fast GLAM (Generalized Linear Array Model) algorithm, and comparable to them in terms of mean squared integrated errors. We extend the sandwich smoother to array data of higher dimensions, where a GLAM algorithm improves the computational speed of the sandwich smoother. One important application of the sandwich smoother is to estimate covariance functions in functional data analysis. In this application, our numerical results show that the sandwich smoother is orders of magnitude faster than local linear regression. The speed of the sandwich formula is important because functional data sets are becoming quite large.

stat.ME

Local Asymptotics of P-splines

This report studies local asymptotics of P-splines with $p$th degree B-splines and a $m$th order difference penalty. Earlier work with $p$ and $m$ restricted is extended to the general case. Asymptotically, penalized splines are kernel estimators with equivalent kernels depending on $m$, but not on $p$. A central limit theorem provides simple expressions for the asymptotic mean and variance. Provided it is fast enough, the divergence rate of the number of knots does not affect the asymptotic distribution. The optimal convergence rate of the penalty parameter is given.

math.ST