SearcharxivSearch

arXiv subjects

Erjia Cui

Publications and source records attributed to Erjia Cui.

16 recordsLinked to original sources

Sparse Longitudinal Functional Principal Component Analysis for Episodic Ambulatory Behavioral Assessments

Accurately monitoring mental fatigue is critical for improving workplace safety and productivity. A recent study examined unobtrusively collected smartphone typing speed as a potential ambulatory proxy assessment of mental fatigue using data from the Intern Health Study (IHS). While population-level average typing speed patterns were found to be consistent with validated measures of mental fatigue, how these trajectories vary across participants and days may inform opportune moments for just-in-time interventions and remains an open question. Treating typing speed trajectories as sparsely observed functional data, we propose a novel sparse longitudinal functional principal component analysis (sparse LFPCA) method for decomposing variability and predicting individual curves. Specifically, sparse data are accommodated by casting covariance estimation as a structured penalized spline regression problem, enabling simultaneous estimation and smoothing of multiple covariance components while borrowing information across locations in the functional domain. Simulations show that sparse LFPCA (1) accurately estimates eigenfunctions and generates reasonable predictions for underlying curves, and (2) achieves similar or superior performance compared to existing alternatives. Our analysis of typing speed data collected from IHS reveals new and interpretable participant- and day-level patterns not captured by previous analyses and can be used to tailor behavioral interventions.

stat.ME

PCA score regression: the art of losing power

The regression of principal component scores (RPCS) on covariates is a widely used analytic approach to detect and test for associations between functional measurements and study participant characteristics. Here we show that: (1) RPCS loses power relative to Function on Scalar Regression (FoSR); (2) the amount of power loss depends on the correlation between the PCs and the true effect; (3) if not corrected for multiplicity, RPCS has inflated $\alpha$-level; and (4) current RPCS methods do not provide valid inference for the true effect. In contrast, we show that Function on Scalar Regression (FoSR) can avoid these problems using a particular combination of modeling tools. We validate these theoretical findings through extensive simulations and illustrate their practical implications using minute-level accelerometry data from the National Health and Nutrition Examination Survey (NHANES).

stat.ME

Quantifying Time-Varying Physical Activity Intervention Effects via Functional Regression

Physical activity (PA) intervention studies often collect repeated intensity measurements over long observation periods. Quantifying the variation in intervention effects over the study period is critical to evaluating and improving intervention strategies, yet many analyses reduce PA data into scalar summary measures, resulting in limited insights. We propose a functional regression framework, which captures time-varying intervention effects by modeling the entire PA trajectory as a functional observation. From both methodological and practical perspectives, we demonstrate the advantages of function-on-scalar regression (FoSR) over the traditional two-step approach of applying functional principal components analysis (FPCA) followed by regressing scores on covariates. The FoSR is further extended to a function-on-function regression (FoFR) for studying the association of PA across time periods. Methods are applied to daily step counts from the Social incentives to Encourage Physical Activity and Understand Predictors (STEP UP) study, revealing distinct and highly interpretable time-varying effects of three intervention strategies on PA and differences in their sustainability. Our case study highlights the feasibility of functional data analysis techniques for uncovering novel insights in intervention studies with high-dimensional endpoints.

stat.AP

BAMIFun: Bayesian Multiple Imputation for Functional Data

Missing data are pervasive in modern functional datasets, where trajectories are often sparsely or irregularly observed. Although Functional Principal Component Analysis (FPCA) is widely used to reconstruct incomplete curves, existing approaches typically employ single imputation, leading to overly optimistic inferences in downstream analyses. To address these challenges, we develop a novel Bayesian multiple imputation framework for functional data (BAMIFun). For single-level functional data, we impose a Bayesian low-rank model that incorporates penalized spline representations to enforce smoothness of the functional domain and derive an efficient Gibbs sampler algorithm for posterior computation. In addition, we demonstrate and validate how to properly account for estimation uncertainties in downstream analysis. Furthermore, we extend the framework to multiway functional data using Functional Tensor Singular Value Decomposition (FTSVD) model, enabling Bayesian multiple imputation in settings not supported by existing methods. Simulation studies show that BAMIFun achieves substantially improved coverage and more reliable downstream inference compared to existing methods, while maintaining similar imputation accuracy. Case studies using a physical activity dataset and an infant gut microbiome dataset further demonstrate the practical advantages of our proposed methods under severe missingness. The code is available at https://github.com/ZirenJiang/BAMIFun.

stat.ME

Efficient Longitudinal Function-on-Function Regression

We propose a computationally efficient inferential procedure for longitudinal function-on-function regression. The method follows a marginal three-step approach: (1) fit massive pointwise longitudinal scalar-on-function regression models, (2) smooth the resulting estimates along the bivariate functional domain, and (3) compute confidence bands using either an analytic approach for Gaussian data or a cluster bootstrap for Gaussian or non-Gaussian data. Simulation studies demonstrate that the proposed method achieves accurate estimation and valid inference, while substantially reducing computational burden compared to existing approaches. Methods are motivated by a physical activity intervention trial in older adults where high-dimensional wearable data were collected longitudinally across multiple visits. Our applications reveal significant increases in physical activity in the morning using interpersonal intervention strategies, but not intrapersonal strategies. The proposed methods are implemented in an R package.

stat.ME

Estimating causal effects of functional treatments with modified functional treatment policies

Functional data are increasingly prevalent in biomedical research. While functional data analysis has been established for decades, causal inference with functional treatments remains largely unexplored. Existing methods typically focus on estimating the causal average dose response functional (ADRF), which requires strong positivity assumptions and offers limited interpretability. In this work, we target a new causal estimand, the modified functional treatment policy (MFTP), which focuses on estimating the average potential outcome when each individual slightly modifies their treatment trajectory from the observed one. A major challenge for this new estimand is the need to define an average over an infinite-dimensional object with no density. By proposing a novel definition of the population average over a functional variable using a functional principal component analysis (FPCA) decomposition, we establish the causal identifiability of the MFTP estimand. We further derive outcome regression, inverse probability weighting, and doubly robust estimators for the MFTP, and provide theoretical guarantees under mild regularity conditions. The proposed estimators are validated through extensive simulation studies. Applying our MFTP framework to the National Health and Nutrition Examination Survey (NHANES) accelerometer data, we estimate the causal effects of reducing disruptive nighttime activity and low-activity duration on all-cause mortality.

stat.ME

Inference for the Extended Functional Cox Model: A UK Biobank Case Study

Multiple studies have shown that scalar summaries of objectively measured physical activity (PA) using accelerometers are the strongest predictors of mortality, outperforming all traditional risk factors, including age, sex, body mass index (BMI), and smoking. Here we show that diurnal patterns of PA and their day-to-day variability provide additional information about mortality. To do that, we introduce a class of extended functional Cox models and corresponding inferential tools designed to quantify the association between multiple functional and scalar predictors with time-to-event outcomes in large-scale (large $n$) high-dimensional (large $p$) datasets. Methods are applied to the UK Biobank study, which collected PA at every minute of the day for up to seven days, as well as time to mortality ($93{,}370$ participants with good quality accelerometry data and $931$ events). Simulation studies show that methods perform well in realistic scenarios and scale up to studies an order of magnitude larger than the UK Biobank accelerometry study. Establishing the feasibility and scalability of these methods for such complex and large data sets is a major milestone in applied Functional Data Analysis (FDA).

stat.ME

Functional Accelerated Failure Time Models for Predicting Time Since Cannabis Use

Cannabis consumption impairs key driving skills and increases crash risk, yet few objective, validated tools exists to identify acute cannabis use or impairment in traffic safety settings. Pupil response to light has emerged as a promising biomarker of recent cannabis use, but its predictive utility remains underexplored. We propose two functional accelerated failure time (AFT) models for predicting time since cannabis use from pupil light response curves. The linear functional AFT (lfAFT) model provides a simple and interpretable framework that summarizes the overall contribution of a functional covariate to time-since-smoking, while the additive functional AFT (afAFT) model generalizes this structure by allowing effects to vary flexibly with both magnitude and location of the functional covariate. Estimation is computationally efficient and straightforward to implement. Simulation studies show that the proposed methods achieve strong estimation accuracy and predictive performance across various scenarios and remain robust to moderate model misspecification. Application to pupillometry data from the Colorado Cannabis & Driving Study demonstrates that pupil light response curves contain meaningful predictive signal, underscoring the potential of these models for traffic safety and broader biomedical applications.

stat.ME

Fast Penalized Generalized Estimating Equations for Large Longitudinal Functional Datasets

Longitudinal binary or count functional data are common in neuroscience, but are often too large to analyze with existing functional regression methods. We propose one-step penalized generalized estimating equations that supports generalized functional outcomes (e.g., count, binary, proportion, continuous-valued) and is fast even when datasets have a large number of clusters and large cluster sizes. The method applies to functional and scalar covariates and the one-step estimation framework enables efficient smoothing parameter selection, bootstrapping, and joint confidence interval construction. Importantly, this semi-parametric approach yields coefficient confidence intervals that are provably valid asymptotically even under working correlation misspecification. By developing a general theory for adaptive one-step M-estimation, we prove that the coefficient estimates are asymptotically normal and as efficient as the fully-iterated estimator; we verify these theoretical properties in simulations. We illustrate the benefits of our approach for analyzing large-scale neural recordings by applying it to a recent calcium imaging dataset published in Nature. We show that our method reveals important timing effects obscured in non-functional analyses. In doing so, we also demonstrate scaling to common neuroscience dataset sizes: the one-step estimator fits to a dataset with 150,000 (binary) functional outcomes, each observed at 120 functional domain points, in only 6.5 minutes on a laptop without parallelization. We release our methods in the R package 'fastfGEE', which supports a wide range of link functions and working covariances.

stat.ME

Tutorial on Bayesian Functional Regression Using Stan

This manuscript provides step-by-step instructions for implementing Bayesian functional regression models using Stan. Extensive simulations indicate that the inferential performance of the methods is comparable to that of state-of-the-art frequentist approaches. However, Bayesian approaches allow for more flexible modeling and provide an alternative when frequentist methods are not available or may require additional development. Methods and software are illustrated using the accelerometry data from the National Health and Nutrition Examination Survey (NHANES).

stat.ME

Prediction Inference Using Generalized Functional Mixed Effects Models

We introduce inferential methods for prediction based on functional random effects in generalized functional mixed effects models. This is similar to the inference for random effects in generalized linear mixed effects models (GLMMs), but for functional instead of scalar outcomes. The method combines: (1) local GLMMs to extract initial estimators of the functional random components on the linear predictor scale; (2) structural functional principal components analysis (SFPCA) for dimension reduction; and (3) global Bayesian multilevel model conditional on the eigenfunctions for inference on the functional random effects. Extensive simulations demonstrate excellent coverage properties of credible intervals for the functional random effects in a variety of scenarios and for different data sizes. To our knowledge, this is the first time such simulations are conducted and reported, likely because prediction inference was not viewed as a priority and existing methods are too slow to calculate coverage. Methods are implemented in a reproducible R package and demonstrated using the NHANES 2011-2014 accelerometry data.

stat.ME

Generalized Conditional Functional Principal Component Analysis

We propose generalized conditional functional principal components analysis (GC-FPCA) for the joint modeling of the fixed and random effects of non-Gaussian functional outcomes. The method scales up to very large functional data sets by estimating the principal components of the covariance matrix on the linear predictor scale conditional on the fixed effects. This is achieved by combining three modeling innovations: (1) fit local generalized linear mixed models (GLMMs) conditional on covariates in windows along the functional domain; (2) conduct a functional principal component analysis (FPCA) on the person-specific functional effects obtained by assembling the estimated random effects from the local GLMMs; and (3) fit a joint functional mixed effects model conditional on covariates and the estimated principal components from the previous step. GC-FPCA was motivated by modeling the minute-level active/inactive profiles over the day ($1{,}440$ 0/1 measurements per person) for $8{,}700$ study participants in the National Health and Nutrition Examination Survey (NHANES) 2011-2014. We show that state-of-the-art approaches cannot handle data of this size and complexity, while GC-FPCA can.

stat.ME

Continuous-time multivariate analysis

The starting point for much of multivariate analysis (MVA) is an $n\times p$ data matrix whose $n$ rows represent observations and whose $p$ columns represent variables. Some multivariate data sets, however, may be best conceptualized not as $n$ discrete $p$-variate observations, but as $p$ curves or functions defined on a common time interval. Here we introduce a framework for extending techniques of multivariate analysis to such settings. The proposed continuous-time multivariate analysis (CTMVA) framework rests on the assumption that the curves can be represented as linear combinations of basis functions such as $B$-splines, as in the Ramsay-Silverman representation of functional data; but whereas functional data analysis extends MVA to the case of observations that are curves rather than vectors -- heuristically, $n\times p$ data with $p$ infinite -- we are instead concerned with what happens when $n$ is infinite. We present continuous-time extensions of the classical MVA methods of covariance and correlation estimation, principal component analysis, Fisher's linear discriminant analysis, and $k$-means clustering. We show that CTMVA can improve on the performance of classical MVA, in particular for correlation estimation and clustering, and can be applied in some settings where classical MVA cannot, including variables observed at disparate time points. CTMVA is illustrated with a novel perspective on a well-known Canadian weather data set, and with applications to data sets involving international development, brain signals, and air quality. The proposed methods are implemented in the publicly available R package \texttt{ctmva}.

stat.ME

Scalable regression calibration approaches to correcting measurement error in multi-level generalized functional linear regression models with heteroscedastic measurement errors

Wearable devices permit the continuous monitoring of biological processes, such as blood glucose metabolism, and behavior, such as sleep quality and physical activity. The continuous monitoring often occurs in epochs of 60 seconds over multiple days, resulting in high dimensional longitudinal curves that are best described and analyzed as functional data. From this perspective, the functional data are smooth, latent functions obtained at discrete time intervals and prone to homoscedastic white noise. However, the assumption of homoscedastic errors might not be appropriate in this setting because the devices collect the data serially. While researchers have previously addressed measurement error in scalar covariates prone to errors, less work has been done on correcting measurement error in high dimensional longitudinal curves prone to heteroscedastic errors. We present two new methods for correcting measurement error in longitudinal functional curves prone to complex measurement error structures in multi-level generalized functional linear regression models. These methods are based on two-stage scalable regression calibration. We assume that the distribution of the scalar responses and the surrogate measures prone to heteroscedastic errors both belong in the exponential family and that the measurement errors follow Gaussian processes. In simulations and sensitivity analyses, we established some finite sample properties of these methods. In our simulations, both regression calibration methods for correcting measurement error performed better than estimators based on averaging the longitudinal functional data and using observations from a single day. We also applied the methods to assess the relationship between physical activity and type 2 diabetes in community dwelling adults in the United States who participated in the National Health and Nutrition Examination Survey.

stat.ME

A case study of glucose levels during sleep using fast function on scalar regression inference

Continuous glucose monitors (CGMs) are increasingly used to measure blood glucose levels and provide information about the treatment and management of diabetes. Our motivating study contains CGM data during sleep for 174 study participants with type II diabetes mellitus measured at a 5-minute frequency for an average of 10 nights. We aim to quantify the effects of diabetes medications and sleep apnea severity on glucose levels. Statistically, this is an inference question about the association between scalar covariates and functional responses. However, many characteristics of the data make analyses difficult, including (1) non-stationary within-day patterns; (2) substantial between-day heterogeneity, non-Gaussianity, and outliers; 3) large dimensionality due to the number of study participants, sleep periods, and time points. We evaluate and compare two methods: fast univariate inference (FUI) and functional additive mixed models (FAMM). We introduce a new approach for calculating p-values for testing a global null effect of covariates using FUI, and provide practical guidelines for speeding up FAMM computations, making it feasible for our data. While FUI and FAMM are philosophically different, they lead to similar point estimators in our study. In contrast to FAMM, FUI is fast, accounts for within-day correlations, and enables the construction of joint confidence intervals. Our analyses reveal that: (1) biguanide medication and sleep apnea severity significantly affect glucose trajectories during sleep, and (2) the estimated effects are time-invariant.

stat.AP

On the distribution of the hitting time for the N-urn Ehrenfest model

In this paper, we consider the N-urn Ehrenfest model. By utilizing an auxiliary continuous-time Markov chain, we obtain the explicit formula for the Laplace transform of the hitting time from a single state to a set A of states where A satisfies some symmetric properties. After obtaining the Laplace transform, we are able to compute the high-order moments(especially, variance) for the hitting time.

math.PR