SearcharxivSearch

arXiv subjects

Vadim Zipunnikov

Publications and source records attributed to Vadim Zipunnikov.

At least 19 recordsLinked to original sources

Multivariate Functional Principal Component Analysis for Mixed-Type mHealth Data: An Application to Mood Disorders

Modern mobile health (mHealth) assessment combines self-reported measures of participants' health experiences with passively collected health behavior data throughout the day. These data are collected across multiple measurement scales, including continuous (physical activity), truncated (pain), ordinal (mood), and binary (daily life events). When indexed by time of day and stacked across assessment domains, these data structures can be treated as multivariate functional data comprising continuous, truncated, ordinal, and binary variables. Motivated by these applications, we propose a multivariate functional principal component analysis for mixed-type data ($M^2$FPCA). The approach is based on a semiparametric Gaussian copula model and assumes that the observed data arise from an underlying multivariate generalized latent nonparanormal functional process. Latent temporal and inter-variable dependence are estimated semiparametrically through Kendall's tau bridging method. Two covariance estimation procedures are developed: a fully multivariate block-wise estimator and a computationally efficient alternative based on partial separability that assumes shared principal components across domains. The proposed method yields interpretable latent functional principal component scores that can serve as participant-specific digital biomarkers. Simulation studies demonstrate the method's competitive performance under various complex dependence structures. The method is applied to mHealth data from 307 participants in the National Institute of Mental Health Family Study of Mood and Affective Spectrum Disorders. Our approach identifies time-of-day patterns shared across mood, anxiety, energy, and physical activity that meaningfully stratify mood disorder subtypes.

stat.ME

Scalar-on-distribution regression via generalized odds with applications to accelerometry-assessed disability in multiple sclerosis

Distributional representations of data collected using digital health technologies have been shown to outperform scalar summaries for clinical prediction, with carefully quantified tail-behavior often driving the gains. Motivated by these findings, we propose a unified generalized odds (GO) framework that represents subject-specific distributions through ratios of probabilities over arbitrary regions of the sample space, subsuming hazard, survival, and residual life representations as special cases. We develop a scale-on-odds regression model using spline-based functional representations with penalization for efficient estimation. Applied to wrist-worn accelerometry data from the HEAL-MS study, generalized odds models yield improved prediction of Expanded Disability Status Scale (EDSS) scores compared to classical scalar and survival-based approaches, demonstrating the value of odds-based distributional covariates for modeling DHT data.

stat.ME

Hazard and Beyond: Exploring Five Distributional Representations of Accelerometry Data for Disability Discrimination in Multiple Sclerosis

Research on modeling the distributional aspects in sensor-based digital health (sDHT) data has grown significantly in recent years. Most existing approaches focus on using individual-specific density or quantile functions. However, there has been limited exploration to assess the practical utility of alternative distributional representations in clinical contexts collecting sDHT data. This study is motivated by accelerometry data collected on 246 individuals with multiple sclerosis (MS)representing a wide range of disability (Expanded Disability Status Scale, EDSS: 0-7). We consider five different individual-level distributional representations of minute-level activity counts: density, survival, hazard, quantile, and total time on test functions. For each of the five distributional representations, scalar-on-function regression fits linear discriminators for binary and continuously measured MS disability, and cross-validated discriminatory performance of these linear discriminators is compared across. The results show that individual-level hazard functions provide the highest discriminatory accuracy, more than double the accuracy compared to density functions. Individual-level quantile functions provided the second-highest discriminatory accuracy. These findings highlight the importance of focusing on distributional representations that capture the tail behavior of distributions when analyzing digital health data, especially in clinical contexts.

stat.AP

Glucodensity Functional Profiles Outperform Traditional Continuous Glucose Monitoring Metrics

Continuous glucose monitoring (CGM) data has revolutionized the management of type 1 diabetes, particularly when integrated with insulin pumps to mitigate clinical events such as hypoglycemia. Recently, there has been growing interest in utilizing CGM devices in clinical studies involving healthy and diabetes populations. However, efficiently exploiting the high temporal resolution of CGM profiles remains a significant challenge. Numerous indices -- such as time-in-range metrics and glucose variability measures -- have been proposed, but evidence suggests these metrics overlook critical aspects of glucose dynamic homeostasis. As an alternative method, this paper explores the clinical value of glucodensity metrics in capturing glucose dynamics -- specifically the speed and acceleration of CGM time series -- as new biomarkers for predicting long-term glucose outcomes. Our results demonstrate significant information gains, exceeding 20\% in terms of adjusted $R^2$, in forecasting glycosylated hemoglobin (HbA1c) and fasting plasma glucose (FPG) at five and eight years from baseline AEGIS data, compared to traditional non-CGM and CGM glucose biomarkers. These findings underscore the importance of incorporating more complex CGM functional metrics, such as the glucodensity approach, to fully capture continuous glucose fluctuations across different time-scale resolutions.

stat.AP

Distributional outcome regression via quantile functions and its application to modelling continuously monitored heart rate and physical activity

Modern clinical and epidemiological studies widely employ wearables to record parallel streams of real-time data on human physiology and behavior. With recent advances in distributional data analysis, these high-frequency data are now often treated as distributional observations resulting in novel regression settings. Motivated by these modelling setups, we develop a distributional outcome regression via quantile functions (DORQF) that expands existing literature with three key contributions: i) handling both scalar and distributional predictors, ii) ensuring jointly monotone regression structure without enforcing monotonicity on individual functional regression coefficients, iii) providing statistical inference via asymptotic projection-based joint confidence bands and a statistical test of global significance to quantify uncertainty of the estimated functional regression coefficients. The method is motivated by and applied to Actiheart component of Baltimore Longitudinal Study of Aging that collected one week of minute-level heart rate (HR) and physical activity (PA) data on 781 older adults to gain deeper understanding of age-related changes in daily life heart rate reserve, defined as a distribution of daily HR, while accounting for daily distribution of physical activity, age, gender, and body composition. Intriguingly, the results provide novel insights in epidemiology of daily life heart rate reserve.

stat.ME

Functional Principal Component Analysis for Continuous non-Gaussian, Truncated, and Discrete Functional Data

Mobile health studies often collect multiple within-day self-reported assessments of participants' behavior and well-being on different scales such as physical activity (continuous), pain levels (truncated), mood states (ordinal), and life events (binary). These assessments, when indexed by time of day, can be treated as functional data of different types - continuous, truncated, ordinal, and binary. We develop a functional principal component analysis that deals with all four types of functional data in a unified manner. It employs a semiparametric Gaussian copula model, assuming a generalized latent non-paranormal process as the underlying mechanism for these four types of functional data. We specify latent temporal dependence using a covariance estimated through Kendall's tau bridging method, incorporating smoothness during the bridging process. Simulation studies demonstrate the method's competitive performance under both dense and sparse sampling conditions. We then apply this approach to data from 497 participants in the National Institute of Mental Health Family Study of the Mood Disorder Spectrum to characterize within-day temporal patterns of mood differences among individuals with major mood disorder subtypes, including Major Depressive Disorder, Type 1, and Type 2 Bipolar Disorder.

stat.ME

Density-on-Density Regression

In this study, a density-on-density regression model is introduced, where the association between densities is elucidated via a warping function. The proposed model has the advantage of a being straightforward demonstration of how one density transforms into another. Using the Riemannian representation of density functions, which is the square-root function (or half density), the model is defined in the correspondingly constructed Riemannian manifold. To estimate the warping function, it is proposed to minimize the average Hellinger distance, which is equivalent to minimizing the average Fisher-Rao distance between densities. An optimization algorithm is introduced by estimating the smooth monotone transformation of the warping function. Asymptotic properties of the proposed estimator are discussed. Simulation studies demonstrate the superior performance of the proposed approach over competing approaches in predicting outcome density functions. Applying to a proteomic-imaging study from the Alzheimer's Disease Neuroimaging Initiative, the proposed approach illustrates the connection between the distribution of protein abundance in the cerebrospinal fluid and the distribution of brain regional volume. Discrepancies among cognitive normal subjects, patients with mild cognitive impairment, and Alzheimer's disease (AD) are identified and the findings are in line with existing knowledge about AD.

stat.ME

Shape-constrained Estimation in Functional Regression with Bernstein Polynomials

Shape restrictions on functional regression coefficients such as non-negativity, monotonicity, convexity or concavity are often available in the form of a prior knowledge or required to maintain a structural consistency in functional regression models. A new estimation method is developed in shape-constrained functional regression models using Bernstein polynomials. Specifically, estimation approaches from nonparametric regression are extended to functional data, properly accounting for shape-constraints in a large class of functional regression models such as scalar-on-function regression (SOFR), function-on-scalar regression (FOSR), and function-on-function regression (FOFR). Theoretical results establish the asymptotic consistency of the constrained estimators under standard regularity conditions. A projection based approach provides point-wise asymptotic confidence intervals for the constrained estimators. A bootstrap test is developed facilitating testing of the shape constraints. Numerical analysis using simulations illustrate improvement in efficiency of the estimators from the use of the proposed method under shape constraints. Two applications include i) modeling a drug effect in a mental health study via shape-restricted FOSR and ii) modeling subject-specific quantile functions of accelerometry-estimated physical activity in the Baltimore Longitudinal Study of Aging (BLSA) as outcomes via shape-restricted quantile-function on scalar regression (QFOSR). R software implementation and illustration of the proposed estimation method and the test is provided.

stat.ME

Regression and Dimension Reduction for Multivariate Mixed-Type Data via Semiparametric Gaussian Copula

Clinical and epidemiological studies encode participant information in multivariate vectors with mixed type variables on continuous, truncated, ordinal, and binary scales. Semiparametric Gaussian Copula (SGC) assumes that observed data is generated by latent multivariate normal random variables which marginals are monotonically transformed and then truncated/ordinalized/binarized. In SGC, the latent correlation matrix fully determines the dependence structure and it is estimated through an inversion of ``bridges'' between Kendall's Tau rank correlations of observed variables and latent correlations. By employing SGC, we develop regression (SGC-Reg), principal component analysis (SGC-PCA), and principal component regression (SGC-PCR) for latent representations of observed data. To build our framework, we make several key contributions: i) establishing novel bridging results for general ordinal type variables, ii) developing regression estimation on the latent space and deriving asymptotic normality of estimators, iii) developing a computationally efficient algorithm that reduces calculation complexity of all steps including calculation of asymptotic covariance matrix from $O(n^4)$ to $O(n\log n)$, iv) developing methods to predict latent representations of observed data and perform imputation of missing data, and v) developing principal component analysis and principal component regression on the latent space. We apply our framework to study the association between a 5-year mortality and 61 frailty-related measures composed of 29 continuous, 17 ordinal, and 15 binary variables in 9478 participants of 1999-2010 waves of National Health and Nutrition Examination Survey (NHANES).

stat.ME

Distributional data analysis via quantile functions and its application to modelling digital biomarkers of gait in Alzheimer's Disease

With the advent of continuous health monitoring with wearable devices, users now generate their unique streams of continuous data such as minute-level step counts or heartbeats. Summarizing these streams via scalar summaries often ignores the distributional nature of wearable data and almost unavoidably leads to the loss of critical information. We propose to capture the distributional nature of wearable data via user-specific quantile functions (QF) and use these QFs as predictors in scalar-on-quantile-function-regression (SOQFR). As an alternative approach, we also propose to represent QFs via user-specific L-moments, robust rank-based analogs of traditional moments, and use L-moments as predictors in SOQFR (SOQFR-L). These two approaches provide two mutually consistent interpretations: in terms of quantile levels by SOQFR and in terms of L-moments by SOQFR-L. We also demonstrate how to deal with multi-modal distributional data via Joint and Individual Variation Explained (JIVE) using L-moments. The proposed methods are illustrated in a study of association of digital gait biomarkers with cognitive function in Alzheimer's disease (AD). Our analysis shows that the proposed methods demonstrate higher predictive performance and attain much stronger associations with clinical cognitive scales compared to simple distributional summaries.

stat.ME

Scalar on time-by-distribution regression and its application for modelling associations between daily-living physical activity and cognitive functions in Alzheimer's Disease

Wearable data is a rich source of information that can provide deeper understanding of links between human behaviours and human health. Existing modelling approaches use wearable data summarized at subject level via scalar summaries using regression techniques, temporal (time-of-day) curves using functional data analysis (FDA), and distributions using distributional data analysis (DDA). We propose to capture temporally local distributional information in wearable data using subject-specific time-by-distribution (TD) data objects. Specifically, we propose scalar on time-by-distribution regression (SOTDR) to model associations between scalar response of interest such as health outcomes or disease status and TD predictors. We show that TD data objects can be parsimoniously represented via a collection of time-varying L-moments that capture distributional changes over the time-of-day. The proposed method is applied to the accelerometry study of mild Alzheimer's disease (AD). Mild AD is found to be significantly associated with reduced maximal level of physical activity, particularly during morning hours. It is also demonstrated that TD predictors attain much stronger associations with clinical cognitive scales of attention, verbal memory, and executive function when compared to predictors summarized via scalar total activity counts, temporal functional curves, and quantile functions. Taken together, the present results suggest that the SOTDR analysis provides novel insights into cognitive function and AD.

stat.AP

Connecting population-level AUC and latent scale-invariant $R^2$ via Semiparametric Gaussian Copula and rank correlations

Area Under the Curve (AUC) is arguably the most popular measure of classification accuracy. We use a semiparametric framework to introduce a latent scale-invariant $R^2$, a novel measure of variation explained for an observed binary outcome and an observed continuous predictor, and then directly link the latent $R^2$ to AUC. This enables a mutually consistent simultaneous use of AUC as a measure of classification accuracy and the latent $R^2$ as a scale-invariant measure of explained variation. Specifically, we employ Semiparametric Gaussian Copula (SGC) to model a joint dependence between observed binary outcome and observed continuous predictor via the correlation of latent standard normal random variables. Under SGC, we show how, both population-level AUC and latent scale-invariant $R^2$, defined as a squared latent correlation, can be estimated using any of the four rank statistics calculated on binary-continuous pairs: Wilcoxon rank-sum, Kendall's Tau, Spearman's Rho, and Quadrant rank correlations. We then focus on three implications and applications: i) we explicitly show that under SGC, the population-level AUC and the population-level latent $R^2$ are related via a monotone function that depends on the population-level prevalence rate, ii) we propose Quadrant rank correlation as a robust semiparametric version of AUC; iii) we demonstrate how, under complex-survey designs, Wilcoxon rank sum statistics and Spearman and Quadrant rank correlations provide asymptotically consistent estimators of the population-level AUC using only single-participant survey weights. We illustrate these applications using binary outcome of five-year mortality and continuous predictors including Albumin, Systolic Blood Pressure, and accelerometry-derived measures of total volume of physical activity collected in 2003-2006 National Health and Nutrition Examination Survey (NHANES) cohorts.

stat.ME

Prediction of sustained harmonic walking in the free-living environment using raw accelerometry data

Objective. Using raw, sub-second level, accelerometry data, we propose and validate a method for identifying and characterizing walking in the free-living environment. We focus on the sustained harmonic walking (SHW), which we define as walking for at least 10 seconds with low variability of step frequency. Approach. We utilize the harmonic nature of SHW and quantify local periodicity of the tri-axial raw accelerometry data. We also estimate fundamental frequency of observed signals and link it to the instantaneous walking (step-to-step) frequency (IWF). Next, we report total time spent in SHW, number and durations of SHW bouts, time of the day when SHW occurred and IWF for 49 healthy, elderly individuals. Main results. Sensitivity of the proposed classification method was found to be 97%, while specificity ranged between 87% and 97% and prediction accuracy between 94% and 97%. We report total time in SHW between 140 and 10 minutes-per-day distributed between 340 and 50 bouts. We estimate the average IWF to be 1.7 steps-per-second. Significance. We propose a simple approach for detection of SHW and estimation of IWF, based on Fourier decomposition. The resulting approach is fast and allows processing of a week-long raw accelerometry data (approx. 150 million measurements) in relatively short time (~half an hour) on a common laptop computer (2.8 GHz Intel Core i7, 16 GB DDR3 RAM).

stat.AP

Validation of gait characteristics extracted from raw accelerometry during walking against measures of physical function, mobility, fatigability, and fitness

Background. Wearable accelerometry devices allow collection of high-density activity data in large epidemiological studies both in-the-lab as well as in-the-wild (free-living). Such data can be used to detect and identify periods of sustained harmonic walking. This report aims to establish whether the micro- and macro-features of walking identified in the laboratory and free-living environments are associated with measures of physical function, mobility, fatigability, and fitness. Methods. Fifty-one older adults (median age 77.5) enrolled in the Developmental Epidemiologic Cohort Study in Pittsburgh, Pennsylvania were included in the analyses. The study included an in-the-lab component as well as 7 days of monitoring in-the-wild. Participants were equipped with hip-worn Actigraph GT3X+ activity monitors, which collect high-density raw accelerometry data. We applied a walking identification algorithm to the data and defined features of walking, such as participant-specific walking acceleration and cadence. The association between these walking features and physical function, mobility, fatigability, and fitness was quantified using linear regression analysis. Results. Micro-scale features of walking (acceleration and cadence) estimated from in-the-lab and in-the-wild data were associated with measures of physical function, mobility, fatigability, and fitness. In-the-lab median walking acceleration was strongly inversely associated with physical function, mobility, fatigability and fitness. Additionally, in-the-wild daily walking time was inversely associated with usual- and fast-paced 400m walking time. Conclusions. The proposed accelerometry-derived walking features are significantly associated with measures of physical function, mobility, fatigability, and fitness, which provides evidence of convergent validity.

q-bio.QM

Epidemiology of Objectively Measured Bedtime and Chronotype in the US adolescents and adults: NHANES 2003-2006

Background: We propose a method for estimating the timing of in-bed intervals using objective data in a large representative U.S. sample, and quantify the association between these intervals and age, sex, and day of the week. Methods: The study included 11,951 participants six years and older from the National Health and Nutrition Examination Survey (NHANES) 2003-2006, who wore accelerometers to measure physical activity for seven consecutive days. Participants were instructed to remove the device just before the nighttime sleep period and put it back on immediately after. This nighttime period of non-wear was defined in this paper as the objective bedtime (OBT), an objectively estimated record of the in-bed-interval. For each night of the week, we estimated two measures: the duration of the OBT (OBT-D) and, as a measure of the chronotype, the midpoint of the OBT (OBT-M). We estimated day-of-the-week-specific OBT-D and OBT-M using gender-specific population percentile curves. Differences in OBT-M (chronotype) and OBT-D (the amount of time spent in bed) by age and sex were estimated using regression models. Results: The estimates of OBT-M and their differences among age groups were consistent with the estimates of chronotype obtained via self-report in European populations. The average OBT-M varied significantly by age, while OBT-D was less variable with age. The most pronounced differences were observed between OBT-M of weekday and weekend nights. Conclusions: The proposed measures, OBT-D and OBT-M, provide useful information of time in bed and chronotype in NHANES 2003-2006. They identify within-week patterns of bedtime and can be used to study associations between the bedtime and the large number of health outcomes collected in NHANES 2003-2006.

stat.AP

Stride variability measures derived from wrist- and hip-worn accelerometers

Many epidemiological and clinical studies use accelerometry to objectively measure physical activity using the activity counts, vector magnitude, or number of steps. These measures use just a fraction of the information in the raw accelerometry data as they are typically summarized at the minute level. To address this problem we define and estimate two gait measures of temporal stride-to-stride variability based on raw accelerometry data: Amplitude Deviation (AD) and Phase Deviation (PD). We explore the sensitivity of our approach to on-body placement of the accelerometer by comparing hip, left and right wrist placements. We illustrate the approach by estimating AD and PD in 46 elderly participants in the Developmental Epidemiologic Cohort Study (DECOS) who worn accelerometers during a 400 meter walk test. We also show that AD and PD have a statistically significant association with the gait speed and sit-to-stand test performance

q-bio.NC

Longitudinal high-dimensional principal components analysis with application to diffusion tensor imaging of multiple sclerosis

We develop a flexible framework for modeling high-dimensional imaging data observed longitudinally. The approach decomposes the observed variability of repeatedly measured high-dimensional observations into three additive components: a subject-specific imaging random intercept that quantifies the cross-sectional variability, a subject-specific imaging slope that quantifies the dynamic irreversible deformation over multiple realizations, and a subject-visit-specific imaging deviation that quantifies exchangeable effects between visits. The proposed method is very fast, scalable to studies including ultrahigh-dimensional data, and can easily be adapted to and executed on modest computing infrastructures. The method is applied to the longitudinal analysis of diffusion tensor imaging (DTI) data of the corpus callosum of multiple sclerosis (MS) subjects. The study includes $176$ subjects observed at $466$ visits. For each subject and visit the study contains a registered DTI scan of the corpus callosum at roughly 30,000 voxels.

stat.AP

Fast, Exact Bootstrap Principal Component Analysis for p>1 million

Many have suggested a bootstrap procedure for estimating the sampling variability of principal component analysis (PCA) results. However, when the number of measurements per subject ($p$) is much larger than the number of subjects ($n$), the challenge of calculating and storing the leading principal components from each bootstrap sample can be computationally infeasible. To address this, we outline methods for fast, exact calculation of bootstrap principal components, eigenvalues, and scores. Our methods leverage the fact that all bootstrap samples occupy the same $n$-dimensional subspace as the original sample. As a result, all bootstrap principal components are limited to the same $n$-dimensional subspace and can be efficiently represented by their low dimensional coordinates in that subspace. Several uncertainty metrics can be computed solely based on the bootstrap distribution of these low dimensional coordinates, without calculating or storing the $p$-dimensional bootstrap components. Fast bootstrap PCA is applied to a dataset of sleep electroencephalogram (EEG) recordings ($p=900$, $n=392$), and to a dataset of brain magnetic resonance images (MRIs) ($p\approx$ 3 million, $n=352$). For the brain MRI dataset, our method allows for standard errors for the first 3 principal components based on 1000 bootstrap samples to be calculated on a standard laptop in 47 minutes, as opposed to approximately 4 days with standard methods.

stat.ME