SearcharxivSearch

arXiv subjects

Loki Natarajan

Publications and source records attributed to Loki Natarajan.

12 recordsLinked to original sources

Sedentary Behavior Classification for Wearable Sensors with a CNN-BiLSTM Model

Accurate detection of sedentary behavior is important for studying health risks related to prolonged sitting, but posture-based classification remains challenging with wearable sensors, especially at the wrist. We study whether a deep learning model trained on hip-worn accelerometer data can transfer to wrist-worn accelerometer data for sitting versus non-sitting classification. We use CHAP, a CNN-BiLSTM model originally developed for hip accelerometers, and evaluate its zero-shot performance on wrist data as well as its adaptation through finetuning with varying amounts of labeled wrist data. Experiments are conducted on the iWatch dataset with ground-truth posture labels derived from wearable cameras. The hip-trained model performs strongly on hip data without retraining, but accuracy drops on wrist data due to sensor placement shift. Finetuning CHAP provides consistent advantages over transformer models trained from scratch. These findings suggest that hip-based pretraining provides a useful starting point for wrist deployment, while highlighting the need for wrist-specific adaptation to handle higher signal variability.

cs.LG

Modeling Physical Activity Change as Smooth Transformations: Temporal and Amplitude Patterns Associated with Physical Function in Older Women

Purpose: To investigate whether longitudinal changes in timing and magnitude of PA are associated with physical function (PF) in older women. Methods: Women from OPACH study with accelerometry at baseline and WHISH study W1 and W2 were included. Minute-level PA counts were averaged and smoothed as diurnal PA curves. Consecutive-visit change was modeled within periods (baseline--W1 and W1--W2) as a Riemannian deformation from earlier to later curves, with two-dimensional initial momenta characterizing timing and magnitude shifts. Multivariate functional principal component analysis (MFPCA) summarized coupled timing-magnitude patterns, and principal component (PC) scores and deformation energy were derived for each participant-period. Linear mixed-effects models related these features to RAND-36 PF, adjusting for baseline PF and covariates. Results: Mean PA deformation in both periods showed downward shifts in PA magnitude and temporal redistribution after 10:00. Top 15 PCs explained at least 90% of variability in both periods. PC1 captured diurnal PA increase/decrease, explaining 22.4% of variability for baseline--W1 and 20.8% for W1--W2. Among participants with complete PF scores and baseline covariates (N=1,157), higher PC1 scores, reflecting relative increase/maintenance of PA across day, were positively associated with PF (P<0.0001). Deformation energy, a metric for overall diurnal pattern change between visits, showed a significant interaction with period for PF (P=0.003), with a larger positive association during W1--W2 than during baseline--W1. Conclusions: In older women, longitudinal changes in diurnal PA accumulation were associated with PF. Riemannian deformation analysis identified clinically interpretable markers of PA pattern change that may capture functional-aging information not represented by conventional PA summaries.

stat.AP

MoCA: Multi-modal Cross-masked Autoencoder for Time Series in Digital Health

Wearable devices enable continuous multi-modal physiological and behavioral monitoring, yet analysis of these data streams faces fundamental challenges including the lack of gold-standard labels and incomplete sensor data. While self-supervised learning approaches have shown promise for addressing these issues, existing multi-modal extensions present opportunities to better leverage the rich temporal and cross-modal correlations inherent in simultaneously recorded wearable sensor data. We propose the Multi-modal Cross-masked Autoencoder (MoCA), a self-supervised learning framework that combines transformer architecture with masked autoencoder (MAE) methodology, using a principled cross-modality masking scheme that explicitly leverages correlation structures between sensor modalities. MoCA demonstrates strong performance boosts across reconstruction and downstream classification tasks on diverse benchmark datasets. We further establish theoretical guarantees by establishing a fundamental connection between multi-modal MAE loss and kernelized canonical correlation analysis through a Reproducing Kernel Hilbert Space framework, providing principled guidance for correlation-aware masking strategy design. Our approach offers a novel solution for leveraging unlabeled multi-modal wearable data while handling missing modalities, with broad applications across digital health domains.

stat.ML

A doubly robust estimator for the Mann Whitney Wilcoxon Rank Sum Test when applied for causal inference in observational studies

The Mann-Whitney-Wilcoxon rank sum test (MWWRST) is a widely used method for comparing two treatment groups in randomized control trials, particularly when dealing with highly skewed data. However, when applied to observational study data, the MWWRST often yields invalid results for causal inference. To address this limitation, Wu et al. (2014) introduced an approach that incorporates inverse probability weighting (IPW) into this rank-based statistics to mitigate confounding effects. Subsequently, Mao (2018), Zhang et al. (2019), and Ai et al. (2020) extended this IPW estimator to develop doubly robust estimators. Nevertheless, each of these approaches has notable limitations. Mao's method imposes stringent assumptions that may not align with real-world study data. Zhang et al.'s (2019) estimators rely on bootstrap inference, which suffers from computational inefficiency and lacks known asymptotic properties. Meanwhile, Ai et al. (2020) primarily focus on testing the null hypothesis of equal distributions between two groups, which is a more stringent assumption that may not be well-suited to the primary practical application of MWWRST. In this paper, we aim to address these limitations by leveraging functional response models (FRM) to develop doubly robust estimators. We demonstrate the performance of our proposed approach using both simulated and real study data.

stat.ME

A linear mixed model approach for measurement error adjustment: applications to sedentary behavior assessment from wearable devices

In recent years, wearable devices have become more common to capture a wide range of health behaviors, especially for physical activity and sedentary behavior. These sensor-based measures are deemed to be objective and thus less prone to self-reported biases, inherent in questionnaire assessments. While this is undoubtedly a major advantage, there can still be measurement errors from the device recordings, which pose serious challenges for conducting statistical analysis and obtaining unbiased risk estimates. There is a vast literature proposing statistical methods for adjusting for measurement errors in self-reported behaviors, such as in dietary intake. However, there is much less research on error correction for sensor-based device measures, especially sedentary behavior. In this paper, we address this gap. Exploiting the excessive multiple-day assessments typically collected when sensor devices are deployed, we propose a two-stage linear mixed effect model (LME) based approach to correct bias caused by measurement errors. We provide theoretical proof of the debiasing process using the Best Linear Unbiased Predictors (BLUP), and use both simulation and real data from a cohort study to demonstrate the performance of the proposed approach while comparing to the naïve plug-in approach that directly uses device measures without appropriately adjusting measurement errors. Our results indicate that employing our easy-to-implement BLUP correction method can greatly reduce biases in disease risk estimates and thus enhance the validity of study findings.

stat.ME

A Riemann Manifold Model Framework for Longitudinal Changes in Physical Activity Patterns

Physical activity (PA) is significantly associated with many health outcomes. The wide usage of wearable accelerometer-based activity trackers in recent years has provided a unique opportunity for in-depth research on PA and its relations with health outcomes and interventions. Past analysis of activity tracker data relies heavily on aggregating minute-level PA records into day-level summary statistics, in which important information of PA temporal/diurnal patterns is lost. In this paper we propose a novel functional data analysis approach based on Riemann manifolds for modeling PA and its longitudinal changes. We model smoothed minute-level PA of a day as one-dimensional Riemann manifolds and longitudinal changes in PA in different visits as deformations between manifolds. The variability in changes of PA among a cohort of subjects is characterized via variability in the deformation. Functional principal component analysis is further adopted to model the deformations and PC scores are used as a proxy in modeling the relation between changes in PA and health outcomes and/or interventions. We conduct comprehensive analyses on data from two clinical trials: Reach for Health (RfH) and Metabolism, Exercise and Nutrition at UCSD (MENU), focusing on the effect of interventions on longitudinal changes in PA patterns and how different modes of changes in PA influence weight loss, respectively. The proposed approach reveals unique modes of changes including overall enhanced PA, boosted morning PA, and shifts of active hours specific to each study cohort. The results bring new insights into the study of longitudinal changes in PA and health and have the potential to facilitate designing of effective health interventions and guidelines.

stat.AP

Inference and Prediction Using Functional Principal Components Analysis: Application to Diabetic Kidney Disease Progression in the Chronic Renal Insufficiency Cohort (CRIC) Study

Repeated longitudinal measurements are commonly used to model long-term disease progression, and timing and number of assessments per patient may vary, leading to irregularly spaced and sparse data. Longitudinal trajectories may exhibit curvilinear patterns, in which mixed linear regression methods may fail to capture true trends in the data. We applied functional principal components analysis to model kidney disease progression via estimated glomerular filtration rate (eGFR) trajectories. In a cohort of 2641 participants with diabetes and up to 15 years of annual follow-up from the Chronic Renal Insufficiency Cohort (CRIC) study, we detected novel dominant modes of variation and patterns of diabetic kidney disease (DKD) progression among subgroups defined by the presence of albuminuria. We conducted inferential permutation tests to assess differences in longitudinal eGFR patterns between groups. To determine whether fitting a full cohort model or separate group-specific models is more optimal for modeling long-term trajectories, we evaluated model fit, using our goodness-of-fit procedure, and future prediction accuracy. Our findings indicated advantages for both modeling approaches in accomplishing different objectives. Beyond DKD, the methods described are applicable to other settings with longitudinally assessed biomarkers as indicators of disease progression. Supplementary materials for this article are available online.

stat.AP

Linear mixed model vs two-stage methods: Developing prognostic models of diabetic kidney disease progression

Identifying prognostic factors for disease progression is a cornerstone of medical research. Repeated assessments of a marker outcome are often used to evaluate disease progression, and the primary research question is to identify factors associated with the longitudinal trajectory of this marker. Our work is motivated by diabetic kidney disease (DKD), where serial measures of estimated glomerular filtration rate (eGFR) are the longitudinal measure of kidney function, and there is notable interest in identifying factors, such as metabolites, that are prognostic for DKD progression. Linear mixed models (LMM) with serial marker outcomes (e.g., eGFR) are a standard approach for prognostic model development, namely by evaluating the time and prognostic factor (e.g., metabolite) interaction. However, two-stage methods that first estimate individual-specific eGFR slopes, and then use these as outcomes in a regression framework with metabolites as predictors are easy to interpret and implement for applied researchers. Herein, we compared the LMM and two-stage methods, in terms of bias and mean squared error via analytic methods and simulations, allowing for irregularly spaced measures and missingness. Our findings provide novel insights into when two-stage methods are suitable longitudinal prognostic modeling alternatives to the LMM. Notably, our findings generalize to other disease studies.

stat.CO

Sedentary Behavior Estimation with Hip-worn Accelerometer Data: Segmentation, Classification and Thresholding

Cohort studies are increasingly using accelerometers for physical activity and sedentary behavior estimation. These devices tend to be less error-prone than self-report, can capture activity throughout the day, and are economical. However, previous methods for estimating sedentary behavior based on hip-worn data are often invalid or suboptimal under free-living situations and subject-to-subject variation. In this paper, we propose a local Markov switching model that takes this situation into account, and introduce a general procedure for posture classification and sedentary behavior analysis that fits the model naturally. Our method features changepoint detection methods in time series and also a two stage classification step that labels data into 3 classes(sitting, standing, stepping). Through a rigorous training-testing paradigm, we showed that our approach achieves > 80% accuracy. In addition, our method is robust and easy to interpret.

cs.LG

BayesTime: Bayesian Functional Principal Components for Sparse Longitudinal Data

Modeling non-linear temporal trajectories is of fundamental interest in many application areas, such as in longitudinal microbiome analysis. Many existing methods focus on estimating mean trajectories, but it is also often of value to assess temporal patterns of individual subjects. Sparse principal components analysis (SFPCA) serves as a useful tool for assessing individual variation in non-linear trajectories; however its application to real data often requires careful model selection criteria and diagnostic tools. Here, we propose a Bayesian approach to SFPCA, which allows users to use the efficient leave-one-out cross-validation (LOO) with Pareto-smoothed importance sampling (PSIS) for model selection, and to utilize the estimated shape parameter from PSIS-LOO and also the posterior predictive checks for graphical model diagnostics. This Bayesian implementation thus enables careful application of SFPCA to a wide range of longitudinal data applications.

stat.ME

Utilizing stability criteria in choosing feature selection methods yields reproducible results in microbiome data

Feature selection is indispensable in microbiome data analysis, but it can be particularly challenging as microbiome data sets are high-dimensional, underdetermined, sparse and compositional. Great efforts have recently been made on developing new methods for feature selection that handle the above data characteristics, but almost all methods were evaluated based on performance of model predictions. However, little attention has been paid to address a fundamental question: how appropriate are those evaluation criteria? Most feature selection methods often control the model fit, but the ability to identify meaningful subsets of features cannot be evaluated simply based on the prediction accuracy. If tiny changes to the training data would lead to large changes in the chosen feature subset, then many of the biological features that an algorithm has found are likely to be a data artifact rather than real biological signal. This crucial need of identifying relevant and reproducible features motivated the reproducibility evaluation criterion such as Stability, which quantifies how robust a method is to perturbations in the data. In our paper, we compare the performance of popular model prediction metric MSE and proposed reproducibility criterion Stability in evaluating four widely used feature selection methods in both simulations and experimental microbiome applications. We conclude that Stability is a preferred feature selection criterion over MSE because it better quantifies the reproducibility of the feature selection method.

q-bio.QM

Statistical tests for the intersection of independent lists of genes: Sensitivity, FDR, and type I error control

Public data repositories have enabled researchers to compare results across multiple genomic studies in order to replicate findings. A common approach is to first rank genes according to an hypothesis of interest within each study. Then, lists of the top-ranked genes within each study are compared across studies. Genes recaptured as highly ranked (usually above some threshold) in multiple studies are considered to be significant. However, this comparison strategy often remains informal, in that type I error and false discovery rate (FDR) are usually uncontrolled. In this paper, we formalize an inferential strategy for this kind of list-intersection discovery test. We show how to compute a $p$-value associated with a "recaptured" set of genes, using a closed-form Poisson approximation to the distribution of the size of the recaptured set. We investigate operating characteristics of the test as a function of the total number of studies considered, the rank threshold within each study, and the number of studies within which a gene must be recaptured to be declared significant. We investigate the trade off between FDR control and expected sensitivity (the expected proportion of true-positive genes identified as significant). We give practical guidance on how to design a bioinformatic list-intersection study with maximal expected sensitivity and prespecified control of type I error (at the set level) and false discovery rate (at the gene level). We show how optimal choice of parameters may depend on particular alternative hypothesis which might hold. We illustrate our methods using prostate cancer gene-expression datasets from the curated Oncomine database, and discuss the effects of dependence between genes on the test.

stat.AP