SearcharxivSearch

arXiv subjects

Ciprian Crainiceanu

Publications and source records attributed to Ciprian Crainiceanu.

At least 19 recordsLinked to original sources

PCA score regression: the art of losing power

The regression of principal component scores (RPCS) on covariates is a widely used analytic approach to detect and test for associations between functional measurements and study participant characteristics. Here we show that: (1) RPCS loses power relative to Function on Scalar Regression (FoSR); (2) the amount of power loss depends on the correlation between the PCs and the true effect; (3) if not corrected for multiplicity, RPCS has inflated $\alpha$-level; and (4) current RPCS methods do not provide valid inference for the true effect. In contrast, we show that Function on Scalar Regression (FoSR) can avoid these problems using a particular combination of modeling tools. We validate these theoretical findings through extensive simulations and illustrate their practical implications using minute-level accelerometry data from the National Health and Nutrition Examination Survey (NHANES).

stat.ME

Sufficient conditions for proper posteriors in fully-Bayesian Functional PCA

In a fully-Bayesian Functional Principal Components Analysis (FPCA) the principal components are treated as unknown infinite-dimensional parameters. By projecting the functional principal components on a rich orthonormal spline basis, we show that orthonormality of the principal components is equivalent to orthonormality of the spline coefficients. A penalty on the integral of the second derivative of the functional principal components can be induced on the spline coefficients, where each function has its own smoothing parameter. Finally, each smoothing parameter is treated as an inverse variance component in the associated mixed effects model. In this work, we demonstrate that no additional conditions are required to ensure that the corresponding smoothing prior, and thus the posterior distribution, is proper. This allows the choice of less informative priors, such that smoothing is driven by the data.

stat.ME

Functional Moments Regression

The Gaussian Process (GP) assumption is often used in functional data analysis. We propose a method to assess departures from the GP assumption, both in terms of the shape of the distribution and its potential dependence on covariates, using a sequence of functional moment regressions. Our methods are inspired by and applied to objectively measured minute-level physical activity data from the National Health and Nutrition Examination Survey (NHANES) 2011-2014 study. In this setting, we find that the GP assumption is not satisfied, quantify the associations between functional moments and covariates, and show that standard data transformations, such as the log transformation, do not resolve the discrepancy between assumptions and reality. We further show that when the effect sizes are moderate, inference on the functional fixed effects is largely unaffected by departures from the GP assumption. However, when effect sizes are small, both inference and prediction of subject-level data can be strongly affected. Extensive simulations support these findings. This pragmatic paper presents new methods for real data analysis, with implications for statistical methodology and for understanding human activity and health.

stat.ME

Function on Scalar Regression with Complex Survey Designs

Large health surveys increasingly collect high-dimensional functional data from wearable devices, and function on scalar regression (FoSR) is often used to quantify the relationship between these functional outcomes and scalar covariates such as age and sex. However, existing methods for FoSR fail to account for complex survey design. We introduce inferential methods for FoSR for studies with complex survey designs. The method combines fast univariate inference (FUI) developed for functional data outcomes and survey sampling inferential methods developed for scalar outcomes. Our approach consists of three steps: (1) fit survey weighted GLMs at each point along the functional domain, (2) smooth coefficients along the functional domain, and (3) use balanced repeated replication (BRR) or the Rao-Wu-Yue-Beaumont (RWYB) bootstrap to obtain pointwise and joint confidence bands for the functional coefficients. The method is motivated by association studies between continuous physical activity data and covariates collected in the National Health and Nutrition Examination Survey (NHANES). A first-of-its-kind analytical simulation study and empirical simulation using the NHANES data demonstrates that our method performs better than existing methods that do not account for the survey structure. Finally, application of the method in NHANES shows the practical implications of accounting for survey structure. The method is implemented in the R package svyfosr.

stat.ME

Bayesian Multivariate Sparse Functional Principal Components Analysis

Functional Principal Components Analysis (FPCA) provides a parsimonious, semi-parametric model for multivariate, sparsely-observed functional data. Frequentist FPCA approaches estimate principal components (PCs) from the data, then condition on these estimates in subsequent analyses. As an alternative, we propose a fully-Bayesian inferential framework for multivariate, sparse functional data (MSFAST) which explicitly models the PCs and incorporates their uncertainty. MSFAST builds upon the FAST approach to FPCA for univariate, densely-observed functional data. Like FAST, MSFAST represents PCs using orthonormal splines and samples the orthonormal spline coefficients using parameter expansion. MSFAST extends FAST to multivariate, sparsely-observed data by (1) standardizing each functional covariate to mitigate poor posterior conditioning due to disparate scales; (2) using a better-suited orthogonal spline basis; (3) updating parameterizations for computational stability; (4) introducing routines that leverage multiple cores and threads to accelerate compute; (5) using a Procrustes-based posterior PC alignment procedure; and (6) providing efficient prediction routines. We evaluate MSFAST alongside existing implementations using simulations. MSFAST produces uniquely valid inferences and accurate estimates, particularly in smaller signal-to-noise regimes. MSFAST is motivated by and applied to a study of child growth, with an accompanying vignette illustrating the implementation step-by-step.

stat.ME

Walking Fingerprinting Using Wrist Accelerometry During Activities of Daily Living in NHANES

We propose a method for identifying individuals based on their continuously monitored wrist-worn accelerometry during activities of daily living. The method consists of three steps: (1) using Adaptive Empirical Pattern Transformation (ADEPT), a highly specific method to identify walking; (2) transforming the accelerometry time series into an image that corresponds to the joint distribution of the time series and its lags; and (3) using the resulting images to construct a person-specific walking fingerprint. The method is applied to 15,000 individuals from the National Health and Nutrition Examination Survey (NHANES) with up to 7 days of wrist accelerometry data collected at 80 Hertz. The resulting dataset contains more than 10 terabytes, is roughly 2 to 3 orders of magnitude larger than previous datasets used for activity recognition, is collected in the free living environment, and does not contain labels for walking periods. Using extensive cross-validation studies, we show that our method is highly predictive and can be successfully extended to a large, heterogeneous sample representative of the U.S. population: in the highest-performing model, the correct participant is in the top 1% of predictions 96% of the time.

stat.AP

Tutorial on Bayesian Functional Regression Using Stan

This manuscript provides step-by-step instructions for implementing Bayesian functional regression models using Stan. Extensive simulations indicate that the inferential performance of the methods is comparable to that of state-of-the-art frequentist approaches. However, Bayesian approaches allow for more flexible modeling and provide an alternative when frequentist methods are not available or may require additional development. Methods and software are illustrated using the accelerometry data from the National Health and Nutrition Examination Survey (NHANES).

stat.ME

Prediction Inference Using Generalized Functional Mixed Effects Models

We introduce inferential methods for prediction based on functional random effects in generalized functional mixed effects models. This is similar to the inference for random effects in generalized linear mixed effects models (GLMMs), but for functional instead of scalar outcomes. The method combines: (1) local GLMMs to extract initial estimators of the functional random components on the linear predictor scale; (2) structural functional principal components analysis (SFPCA) for dimension reduction; and (3) global Bayesian multilevel model conditional on the eigenfunctions for inference on the functional random effects. Extensive simulations demonstrate excellent coverage properties of credible intervals for the functional random effects in a variety of scenarios and for different data sizes. To our knowledge, this is the first time such simulations are conducted and reported, likely because prediction inference was not viewed as a priority and existing methods are too slow to calculate coverage. Methods are implemented in a reproducible R package and demonstrated using the NHANES 2011-2014 accelerometry data.

stat.ME

Fast Bayesian Functional Principal Components Analysis

Functional Principal Components Analysis (FPCA) is a widely used analytic tool for dimension reduction of functional data. Traditional implementations of FPCA estimate the principal components from the data, then treat these estimates as fixed in subsequent analyses. To account for the uncertainty of PC estimates, we propose FAST, a fully-Bayesian FPCA with three core components: (1) projection of eigenfunctions onto an orthonormal spline basis; (2) efficient sampling of the orthonormal spline coefficient matrix using a parameter expansion scheme based on polar decomposition; and (3) ordering eigenvalues during sampling. Extensive simulation studies show that FAST is very stable and performs better compared to existing methods. FAST is motivated by and applied to a study of the variability in mealtime glucose from the Dietary Approaches to Stop Hypertension for Diabetes Continuous Glucose Monitoring (DASH4D CGM) study. All relevant STAN code and simulation routines are available as supplementary material.

stat.ME

Comparing Step Counting Algorithms for High-Resolution Wrist Accelerometry Data in NHANES 2011-2014

Purpose: To quantify the relative performance of step counting algorithms in studies that collect free-living high-resolution wrist accelerometry data and to highlight the implications of using these algorithms in translational research. Methods: Five step counting algorithms (four open source and one proprietary) were applied to the publicly available, free-living, high-resolution wrist accelerometry data collected by the National Health and Nutrition Examination Survey (NHANES) in 2011-2014. The mean daily total step counts were compared in terms of correlation, predictive performance, and estimated hazard ratios of mortality. Results: The estimated number of steps were highly correlated (median=0.91, range 0.77 to 0.98), had high and comparable predictive performance of mortality (median concordance=0.72, range 0.70 to 0.73). The distributions of the number of steps in the population varied widely (mean step counts range from 2,453 to 12,169). Hazard ratios of mortality associated with a 500-step increase per day varied among step counting algorithms between HR=0.88 and 0.96, corresponding to a 300% difference in mortality risk reduction ([1-0.88]/[1-0.96]=3). Conclusion: Different step counting algorithms provide correlated step estimates and have similar predictive performance that is better than traditional predictors of mortality. However, they provide widely different distributions of step counts and estimated reductions in mortality risk for a 500-step increase.

stat.AP

Sensitivity analysis for matching on high-dimensional predictors: A case study of racial disparity in US mortality

Matching on a low dimensional vector of scalar covariates consists of constructing groups of individuals in which each individual in a group is within a pre-specified distance from an individual in another group. However, matching in high dimensional spaces is more challenging because the distance can be sensitive to implementation details, caliper width, and measurement error of observations. To partially address these problems, we propose to use extensive sensitivity analyses and identify the main sources of variation and bias. We illustrate these concepts by examining the racial disparity in all-cause mortality in the US using the National Health and Nutrition Examination Survey (NHANES 2003-2006). In particular, we match African Americans to Caucasian Americans on age, gender, BMI and objectively measured physical activity (PA). PA is measured every minute using accelerometers for up to seven days and then transformed into an empirical distribution of all of the minute-level observations. The Wasserstein metric is used as the measure of distance between these participant-specific distributions.

stat.AP

Walking fingerprinting

We consider the problem of predicting an individual's identity from accelerometry data collected during walking. In a previous paper we introduced an approach that transforms the accelerometry time series into an image by constructing its complete empirical autocorrelation distribution. Predictors derived by partitioning this image into grid cells were used in logistic regression to predict individuals. Here we: (1) implement machine learning methods for prediction using the grid cell-derived predictors; (2) derive inferential methods to screen for the most predictive grid cells; and (3) develop a novel multivariate functional regression model that avoids partitioning of the predictor space into cells. Prediction methods are compared on two open source data sets: (1) accelerometry data collected from $32$ individuals walking on a $1.06$ kilometer path; and (2) accelerometry data collected from six repetitions of walking on a $20$ meter path on two separate occasions at least one week apart for $153$ study participants. In the $32$-individual study, all methods achieve at least $95$% rank-1 accuracy, while in the $153$-individual study, accuracy varies from $41$% to $98$%, depending on the method and prediction task. Methods provide insights into why some individuals are easier to predict than others.

stat.AP

Fast Generalized Functional Principal Components Analysis

We propose a new fast generalized functional principal components analysis (fast-GFPCA) algorithm for dimension reduction of non-Gaussian functional data. The method consists of: (1) binning the data within the functional domain; (2) fitting local random intercept generalized linear mixed models in every bin to obtain the initial estimates of the person-specific functional linear predictors; (3) using fast functional principal component analysis to smooth the linear predictors and obtain their eigenfunctions; and (4) estimating the global model conditional on the eigenfunctions of the linear predictors. An extensive simulation study shows that fast-GFPCA performs as well or better than existing state-of-the-art approaches, it is orders of magnitude faster than existing general purpose GFPCA methods, and scales up well with both the number of observed curves and observations per curve. Methods were motivated by and applied to a study of active/inactive physical activity profiles obtained from wearable accelerometers in the NHANES 2011-2014 study. The method can be implemented by any user familiar with mixed model software, though the R package fastGFPCA is provided for convenience.

stat.ME

A case study of glucose levels during sleep using fast function on scalar regression inference

Continuous glucose monitors (CGMs) are increasingly used to measure blood glucose levels and provide information about the treatment and management of diabetes. Our motivating study contains CGM data during sleep for 174 study participants with type II diabetes mellitus measured at a 5-minute frequency for an average of 10 nights. We aim to quantify the effects of diabetes medications and sleep apnea severity on glucose levels. Statistically, this is an inference question about the association between scalar covariates and functional responses. However, many characteristics of the data make analyses difficult, including (1) non-stationary within-day patterns; (2) substantial between-day heterogeneity, non-Gaussianity, and outliers; 3) large dimensionality due to the number of study participants, sleep periods, and time points. We evaluate and compare two methods: fast univariate inference (FUI) and functional additive mixed models (FAMM). We introduce a new approach for calculating p-values for testing a global null effect of covariates using FUI, and provide practical guidelines for speeding up FAMM computations, making it feasible for our data. While FUI and FAMM are philosophically different, they lead to similar point estimators in our study. In contrast to FAMM, FUI is fast, accounts for within-day correlations, and enables the construction of joint confidence intervals. Our analyses reveal that: (1) biguanide medication and sleep apnea severity significantly affect glucose trajectories during sleep, and (2) the estimated effects are time-invariant.

stat.AP

Prediction of sustained harmonic walking in the free-living environment using raw accelerometry data

Objective. Using raw, sub-second level, accelerometry data, we propose and validate a method for identifying and characterizing walking in the free-living environment. We focus on the sustained harmonic walking (SHW), which we define as walking for at least 10 seconds with low variability of step frequency. Approach. We utilize the harmonic nature of SHW and quantify local periodicity of the tri-axial raw accelerometry data. We also estimate fundamental frequency of observed signals and link it to the instantaneous walking (step-to-step) frequency (IWF). Next, we report total time spent in SHW, number and durations of SHW bouts, time of the day when SHW occurred and IWF for 49 healthy, elderly individuals. Main results. Sensitivity of the proposed classification method was found to be 97%, while specificity ranged between 87% and 97% and prediction accuracy between 94% and 97%. We report total time in SHW between 140 and 10 minutes-per-day distributed between 340 and 50 bouts. We estimate the average IWF to be 1.7 steps-per-second. Significance. We propose a simple approach for detection of SHW and estimation of IWF, based on Fourier decomposition. The resulting approach is fast and allows processing of a week-long raw accelerometry data (approx. 150 million measurements) in relatively short time (~half an hour) on a common laptop computer (2.8 GHz Intel Core i7, 16 GB DDR3 RAM).

stat.AP

Validation of gait characteristics extracted from raw accelerometry during walking against measures of physical function, mobility, fatigability, and fitness

Background. Wearable accelerometry devices allow collection of high-density activity data in large epidemiological studies both in-the-lab as well as in-the-wild (free-living). Such data can be used to detect and identify periods of sustained harmonic walking. This report aims to establish whether the micro- and macro-features of walking identified in the laboratory and free-living environments are associated with measures of physical function, mobility, fatigability, and fitness. Methods. Fifty-one older adults (median age 77.5) enrolled in the Developmental Epidemiologic Cohort Study in Pittsburgh, Pennsylvania were included in the analyses. The study included an in-the-lab component as well as 7 days of monitoring in-the-wild. Participants were equipped with hip-worn Actigraph GT3X+ activity monitors, which collect high-density raw accelerometry data. We applied a walking identification algorithm to the data and defined features of walking, such as participant-specific walking acceleration and cadence. The association between these walking features and physical function, mobility, fatigability, and fitness was quantified using linear regression analysis. Results. Micro-scale features of walking (acceleration and cadence) estimated from in-the-lab and in-the-wild data were associated with measures of physical function, mobility, fatigability, and fitness. In-the-lab median walking acceleration was strongly inversely associated with physical function, mobility, fatigability and fitness. Additionally, in-the-wild daily walking time was inversely associated with usual- and fast-paced 400m walking time. Conclusions. The proposed accelerometry-derived walking features are significantly associated with measures of physical function, mobility, fatigability, and fitness, which provides evidence of convergent validity.

q-bio.QM

Epidemiology of Objectively Measured Bedtime and Chronotype in the US adolescents and adults: NHANES 2003-2006

Background: We propose a method for estimating the timing of in-bed intervals using objective data in a large representative U.S. sample, and quantify the association between these intervals and age, sex, and day of the week. Methods: The study included 11,951 participants six years and older from the National Health and Nutrition Examination Survey (NHANES) 2003-2006, who wore accelerometers to measure physical activity for seven consecutive days. Participants were instructed to remove the device just before the nighttime sleep period and put it back on immediately after. This nighttime period of non-wear was defined in this paper as the objective bedtime (OBT), an objectively estimated record of the in-bed-interval. For each night of the week, we estimated two measures: the duration of the OBT (OBT-D) and, as a measure of the chronotype, the midpoint of the OBT (OBT-M). We estimated day-of-the-week-specific OBT-D and OBT-M using gender-specific population percentile curves. Differences in OBT-M (chronotype) and OBT-D (the amount of time spent in bed) by age and sex were estimated using regression models. Results: The estimates of OBT-M and their differences among age groups were consistent with the estimates of chronotype obtained via self-report in European populations. The average OBT-M varied significantly by age, while OBT-D was less variable with age. The most pronounced differences were observed between OBT-M of weekday and weekend nights. Conclusions: The proposed measures, OBT-D and OBT-M, provide useful information of time in bed and chronotype in NHANES 2003-2006. They identify within-week patterns of bedtime and can be used to study associations between the bedtime and the large number of health outcomes collected in NHANES 2003-2006.

stat.AP

Simple fixed-effects inference for complex functional models

We propose simple inferential approaches for the fixed effects in complex functional mixed effects models. We estimate the fixed effects under the independence of functional residuals assumption and then bootstrap independent units (e.g. subjects) to estimate the variability of and conduct inference in the form of hypothesis testing on the fixed effects parameters. Simulations show excellent coverage probability of the confidence intervals and size of tests. Methods are motivated by and applied to the Baltimore Longitudinal Study of Aging (BLSA), though they are applicable to other studies that collect correlated functional data.

stat.ME