SearcharxivSearch

arXiv subjects

Ciprian M. Crainiceanu

Publications and source records attributed to Ciprian M. Crainiceanu.

15 recordsLinked to original sources

Tutorial for Bayesian Factor Models

Bayesian Factor Models (BFM) are well-established models that decompose the observed variability in a set of mean-zero, independent, and uncorrelated factors (random effects). While Factor Analysis (FA) was introduced in 1904 by Spearman, there has been renewed interest in inferential and computational methods that can adapt to large and complex modern data sets that are now routinely collected in a variety of applications. We provide reproducible, harmonized, and fast software for a variety of recent BFMs that allows the direct comparison of methods and provides a one-stop tutorial for the BFMs and their implementation. We neither endorse nor recommend any of the methods for a particular application; we simply provide a previously unavailable harmonized and reproducible common platform for BFMs. The accompanying factorverse R package is available at https://github.com/peterdunson/factorverse.

stat.CO

FLAME: A Model for Duration-Dependent Risk Accumulation in Episodic Temporal Exposures

Emerging technologies enable continuous monitoring of temporal exposures to disease risk factors, leading to complex exposure processes characterized by subject-specific numbers and durations of exposure episodes. A key scientific question is how the number and duration of such episodes influence disease risk. Existing methods typically rely on scalar summaries or time-indexed representations and are not naturally suited to model duration-dependent risk accumulation at the episode level. We introduce the FLexible Accumulation ModEl (FLAME), a semiparametric model for risk accumulation at the level of individual exposure episodes, with duration as the primary driver of risk. FLAME is motivated by and applied to quantifying the association between the duration of intraoperative hypotension and acute kidney injury (AKI) following cardiac surgery. The estimated risk accumulation function reveals that, although 60 one-minute hypotensive episodes are associated with an AKI probability of 0.24, a single sustained 60-minute episode increases that probability to 0.33, representing a 38% increase despite identical total duration. These findings provide actionable insights for intraoperative hemodynamic management and demonstrate the importance of accounting for episodic exposure patterns. While motivated by cardiac surgery, FLAME is broadly applicable to other settings involving high-resolution temporal exposures. An R package, flameRisk, is provided to facilitate application of the method in practice.

stat.ME

Inference for the Extended Functional Cox Model: A UK Biobank Case Study

Multiple studies have shown that scalar summaries of objectively measured physical activity (PA) using accelerometers are the strongest predictors of mortality, outperforming all traditional risk factors, including age, sex, body mass index (BMI), and smoking. Here we show that diurnal patterns of PA and their day-to-day variability provide additional information about mortality. To do that, we introduce a class of extended functional Cox models and corresponding inferential tools designed to quantify the association between multiple functional and scalar predictors with time-to-event outcomes in large-scale (large $n$) high-dimensional (large $p$) datasets. Methods are applied to the UK Biobank study, which collected PA at every minute of the day for up to seven days, as well as time to mortality ($93{,}370$ participants with good quality accelerometry data and $931$ events). Simulation studies show that methods perform well in realistic scenarios and scale up to studies an order of magnitude larger than the UK Biobank accelerometry study. Establishing the feasibility and scalability of these methods for such complex and large data sets is a major milestone in applied Functional Data Analysis (FDA).

stat.ME

Multilevel functional distributional models with application to continuous glucose monitoring in diabetes clinical trials

Continuous glucose monitoring (CGM) is a minimally invasive technology that measures blood glucose every few minutes for weeks or months at a time. CGM data are often collected in the free-living environment and is strongly related to sleep, physical activity and meal intake. As the timing of these activities varies substantially within- and between-individuals, it is difficult to model CGM trajectories as a function of time of day. Therefore, in practice, CGM trajectories are often reduced to one or two scalar summaries of the thousands of measurements collected for a study participant. To alleviate the potential loss of information, the cumulative distribution function (cdf) of the CGM time series was proposed as an alternative. Here we address the problem of conducting inference on cdfs in clinical trials with long follow up and frequent measurements. Our approach provides three major innovations: (1) modeling the entire cdf and preserving its monotonicity; (2) accounting for the cdfs correlation (because they are measured on the same individual), continuity (results are robust to the choice of the probability grid), and differential error (e.g., medians have lower variability than $0.99$ quantiles); and (3) preserving the family-wise error when the observed data are longitudinal samples of cdfs. We focus on modeling data collected by The Juvenile Diabetes Research Foundation Continuous Glucose Monitoring Group in a large clinical trial that collected CGM data every few minutes for 26 weeks. Our basic observation unit is the distribution of CGM observations in a four-week interval. The scientific goals are to: (1) identify and quantify the effects of factors that affect glycaemic control in type 1 diabetes patients (T1D); and (2) identify and characterize the patients who respond to treatment.

stat.ME

Generalized Multilevel Functional Principal Component Analysis with Application to NHANES Active Inactive Patterns

Between 2011 and 2014 NHANES collected objectively measured physical activity data using wrist-worn accelerometers for tens of thousands of individuals for up to seven days. In this study, we analyze minute-level indicators of being active, which can be viewed as binary (since each minute is either active or inactive), multilevel (because there are multiple days of data for each participant), and functional data (because the within-day measurements can be viewed as a function of time). To identify both within- and between-participant directions of variation in these data, we introduce Generalized Multilevel Functional Principal Component Analysis (GM-FPCA), an approach based on the dimension reduction of the linear predictor. Our results indicate that specific activity patterns captured by GM-FPCA are strongly associated with mortality risk. Extensive simulation studies demonstrate that GM-FPCA accurately estimates model parameters, is computationally stable, and scales up with the number of study participants, visits, and observations per visit. R code for implementing the method is provided.

stat.ME

Generalized Conditional Functional Principal Component Analysis

We propose generalized conditional functional principal components analysis (GC-FPCA) for the joint modeling of the fixed and random effects of non-Gaussian functional outcomes. The method scales up to very large functional data sets by estimating the principal components of the covariance matrix on the linear predictor scale conditional on the fixed effects. This is achieved by combining three modeling innovations: (1) fit local generalized linear mixed models (GLMMs) conditional on covariates in windows along the functional domain; (2) conduct a functional principal component analysis (FPCA) on the person-specific functional effects obtained by assembling the estimated random effects from the local GLMMs; and (3) fit a joint functional mixed effects model conditional on covariates and the estimated principal components from the previous step. GC-FPCA was motivated by modeling the minute-level active/inactive profiles over the day ($1{,}440$ 0/1 measurements per person) for $8{,}700$ study participants in the National Health and Nutrition Examination Survey (NHANES) 2011-2014. We show that state-of-the-art approaches cannot handle data of this size and complexity, while GC-FPCA can.

stat.ME

Fast Covariance Estimation for Sparse Functional Data

Smoothing of noisy sample covariances is an important component in functional data analysis. We propose a novel covariance smoothing method based on penalized splines and associated software. The proposed method is a bivariate spline smoother that is designed for covariance smoothing and can be used for sparse functional or longitudinal data. We propose a fast algorithm for covariance smoothing using leave-one-subject-out cross validation. Our simulations show that the proposed method compares favorably against several commonly used methods. The method is applied to a study of child growth led by one of coauthors and to a public dataset of longitudinal CD4 counts.

stat.ME

High-dimensional Multivariate Mediation: with Application to Neuroimaging Data

Mediation analysis has become an important tool in the behavioral sciences for investigating the role of intermediate variables that lie in the path between a randomized treatment and an outcome variable. The influence of the intermediate variable on the outcome is often explored using structural equation models (SEMs), with model coefficients interpreted as possible effects. While there has been significant research on the topic in recent years, little work has been done on mediation analysis when the intermediate variable (mediator) is a high-dimensional vector. In this work we present a new method for exploratory mediation analysis in this setting called the directions of mediation (DMs). The first DM is defined as the linear combination of the elements of a high-dimensional vector of potential mediators that maximizes the likelihood of the SEM. The subsequent DMs are defined as linear combinations of the elements of the high-dimensional vector that are orthonormal to the previous DMs and maximize the likelihood of the SEM. We provide an estimation algorithm and establish the asymptotic properties of the obtained estimators. This method is well suited for cases when many potential mediators are measured. Examples of high-dimensional potential mediators are brain images composed of hundreds of thousands of voxels, genetic variation measured at millions of SNPs, or vectors of thousands of variables in large-scale epidemiological studies. We demonstrate the method using a functional magnetic resonance imaging (fMRI) study of thermal pain where we are interested in determining which brain locations mediate the relationship between the application of a thermal stimulus and self-reported pain.

stat.ME

Improving Reliability of Subject-Level Resting-State fMRI Parcellation with Shrinkage Estimators

A recent interest in resting state functional magnetic resonance imaging (rsfMRI) lies in subdividing the human brain into anatomically and functionally distinct regions of interest. For example, brain parcellation is often used for defining the network nodes in connectivity studies. While inference has traditionally been performed on group-level data, there is a growing interest in parcellating single subject data. However, this is difficult due to the low signal-to-noise ratio of rsfMRI data, combined with typically short scan lengths. A large number of brain parcellation approaches employ clustering, which begins with a measure of similarity or distance between voxels. The goal of this work is to improve the reproducibility of single-subject parcellation using shrinkage estimators of such measures, allowing the noisy subject-specific estimator to "borrow strength" in a principled manner from a larger population of subjects. We present several empirical Bayes shrinkage estimators and outline methods for shrinkage when multiple scans are not available for each subject. We perform shrinkage on raw intervoxel correlation estimates and use both raw and shrinkage estimates to produce parcellations by performing clustering on the voxels. Our proposed method is agnostic to the choice of clustering method and can be used as a pre-processing step for any clustering algorithm. Using two datasets---a simulated dataset where the true parcellation is known and is subject-specific and a test-retest dataset consisting of two 7-minute rsfMRI scans from 20 subjects---we show that parcellations produced from shrinkage correlation estimates have higher reliability and validity than those produced from raw estimates. Application to test-retest data shows that using shrinkage estimators increases the reproducibility of subject-specific parcellations of the motor cortex by up to 30%.

stat.ME

Relating multi-sequence longitudinal intensity profiles and clinical covariates in new multiple sclerosis lesions

Structural magnetic resonance imaging (MRI) can be used to detect lesions in the brains of multiple sclerosis (MS) patients. The formation of these lesions is a complex process involving inflammation, tissue damage, and tissue repair, all of which are visible on MRI. Here we characterize the lesion formation process on longitudinal, multi-sequence structural MRI from 34 MS patients and relate the longitudinal changes we observe within lesions to therapeutic interventions. In this article, we first outline a pipeline to extract voxel level, multi-sequence longitudinal profiles from four MRI sequences within lesion tissue. We then propose two models to relate clinical covariates to the longitudinal profiles. The first model is a principal component analysis (PCA) regression model, which collapses the information from all four profiles into a scalar value. We find that the score on the first PC identifies areas of slow, long-term intensity changes within the lesion at a voxel level, as validated by two experienced clinicians, a neuroradiologist and a neurologist. On a quality scale of 1 to 4 (4 being the highest) the neuroradiologist gave the score on the first PC a median rating of 4 (95% CI: [4,4]), and the neurologist gave it a median rating of 3 (95% CI: [3,3]). In the PCA regression model, we find that treatment with disease modifying therapies (p-value < 0.01), steroids (p-value < 0.01), and being closer to the boundary of abnormal signal intensity (p-value < 0.01) are associated with a return of a voxel to intensity values closer to that of normal-appearing tissue. The second model is a function-on-scalar regression, which allows for assessment of the individual time points at which the covariates are associated with the profiles. In the function-on-scalar regression both age and distance to the boundary were found to have a statistically significant association with the profiles.

stat.AP

Longitudinal high-dimensional principal components analysis with application to diffusion tensor imaging of multiple sclerosis

We develop a flexible framework for modeling high-dimensional imaging data observed longitudinally. The approach decomposes the observed variability of repeatedly measured high-dimensional observations into three additive components: a subject-specific imaging random intercept that quantifies the cross-sectional variability, a subject-specific imaging slope that quantifies the dynamic irreversible deformation over multiple realizations, and a subject-visit-specific imaging deviation that quantifies exchangeable effects between visits. The proposed method is very fast, scalable to studies including ultrahigh-dimensional data, and can easily be adapted to and executed on modest computing infrastructures. The method is applied to the longitudinal analysis of diffusion tensor imaging (DTI) data of the corpus callosum of multiple sclerosis (MS) subjects. The study includes $176$ subjects observed at $466$ visits. For each subject and visit the study contains a registered DTI scan of the corpus callosum at roughly 30,000 voxels.

stat.AP

Movement Prediction Using Accelerometers in a Human Population

We introduce statistical methods for predicting the types of human activity at sub-second resolution using triaxial accelerometry data. The major innovation is that we use labeled activity data from some subjects to predict the activity labels of other subjects. To achieve this, we normalize the data across subjects by matching the standing up and lying down portions of triaxial accelerometry data. This is necessary to account for differences between the variability in the position of the device relative to gravity, which are induced by body shape and size as well as by the ambiguous definition of device placement. We also normalize the data at the device level to ensure that the magnitude of the signal at rest is similar across devices. After normalization we use overlapping movelets (segments of triaxial accelerometry time series) extracted from some of the subjects to predict the movement type of the other subjects. The problem was motivated by and is applied to a laboratory study of 20 older participants who performed different activities while wearing accelerometers at the hip. Prediction results based on other people's labeled dictionaries of activity performed almost as well as those obtained using their own labeled dictionaries. These findings indicate that prediction of activity types for data collected during natural activities of daily living may actually be possible.

stat.AP

Soft Null Hypotheses: A Case Study of Image Enhancement Detection in Brain Lesions

This work is motivated by a study of a population of multiple sclerosis (MS) patients using dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) to identify active brain lesions. At each visit, a contrast agent is administered intravenously to a subject and a series of images is acquired to reveal the location and activity of MS lesions within the brain. Our goal is to identify and quantify lesion enhancement location at the subject level and lesion enhancement patterns at the population level. With this example, we aim to address the difficult problem of transforming a qualitative scientific null hypothesis, such as "this voxel does not enhance", to a well-defined and numerically testable null hypothesis based on existing data. We call the procedure "soft null hypothesis" testing as opposed to the standard "hard null hypothesis" testing. This problem is fundamentally different from: 1) testing when a quantitative null hypothesis is given; 2) clustering using a mixture distribution; or 3) identifying a reasonable threshold with a parametric null assumption. We analyze a total of 20 subjects scanned at 63 visits (~30Gb), the largest population of such clinical brain images.

stat.AP

Structured Functional Principal Component Analysis

Motivated by modern observational studies, we introduce a class of functional models that expands nested and crossed designs. These models account for the natural inheritance of correlation structure from sampling design in studies where the fundamental sampling unit is a function or image. Inference is based on functional quadratics and their relationship with the underlying covariance structure of the latent processes. A computationally fast and scalable estimation procedure is developed for ultra-high dimensional data. Methods are illustrated in three examples: high-frequency accelerometer data for daily activity, pitch linguistic data for phonetic analysis, and EEG data for studying electrical brain activity during sleep.

stat.AP

Multilevel functional principal component analysis

The Sleep Heart Health Study (SHHS) is a comprehensive landmark study of sleep and its impacts on health outcomes. A primary metric of the SHHS is the in-home polysomnogram, which includes two electroencephalographic (EEG) channels for each subject, at two visits. The volume and importance of this data presents enormous challenges for analysis. To address these challenges, we introduce multilevel functional principal component analysis (MFPCA), a novel statistical methodology designed to extract core intra- and inter-subject geometric components of multilevel functional data. Though motivated by the SHHS, the proposed methodology is generally applicable, with potential relevance to many modern scientific studies of hierarchical or longitudinal functional outcomes. Notably, using MFPCA, we identify and quantify associations between EEG activity during sleep and adverse cardiovascular outcomes.

stat.AP