Searcharxiv⌕ Search

arXiv subjects

Sang Kyu Lee

Publications and source records attributed to Sang Kyu Lee.

5 recordsLinked to original sources

Bayesian Nonparametric Causal Inference for Quantile Residual Life: An Application to Alzheimer's Disease

In Alzheimer's disease research, a clinically important question is how much longer individuals would remain dementia-free beyond a given time under different baseline amyloid statuses. We address this question using observational data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) treating baseline amyloid status as the exposure. Estimation is challenging because amyloid status is confounded, time to dementia onset is heterogeneous and heavily right censored, and the target population depends on joint potential event times. At each time point, we consider the always-survivor principal stratum comprising individuals who would remain dementia-free under both amyloid status and estimate quantile contrasts in residual time to dementia onset. We model the joint distribution of event time, exposure, and baseline covariates using an enriched Dirichlet process mixture and conduct posterior inference via Bayesian g-computation. The framework accommodates partially observed covariates under a within-subcluster missing-at-random assumption, estimates contrasts across multiple time points and quantiles from one posterior fit and supports sensitivity analyses for unmeasured confounding, cross-world dependence, and informative censoring. Simulations show favorable finite-sample performance under heterogeneity and heavy censoring. In ADNI, residual time to dementia onset was shorter under elevated than non-elevated baseline amyloid status, both overall and within baseline subgroups.

stat.AP↗

High-dimensional partial linear model with trend filtering

Understanding the links between diet, metabolic changes, and health outcomes is a key focus in nutritional science and broader biological research. Analyzing relationships, such as those between ultra-processed food (UPF) intake and metabolites, offers insights into potential biomarkers for diet-related diseases and public health applications. However, these analyses are challenging due to high-dimensional data structures and complex, often nonlinear associations between covariates and health outcomes. Traditional linear models and conventional nonparametric methods often lack the flexibility to accurately capture such complexities in biological data. To address these challenges, we propose a high-dimensional partial linear regression model that captures both linear and nonlinear effects, combining the interpretability of linear models with the adaptability of nonparametric approaches. Our model leverages trend filtering to handle local smoothness variations effectively and achieves minimax optimal rates, making it suitable for complex biological datasets. We apply this model to data from the Interactive Diet and Activity Tracking in AARP (IDATA) Study, demonstrating its utility in identifying biomarkers associated with UPF intake and illustrating its potential for broader applications in dietary, metabolic, and health-related research.

stat.ME↗

A Causal Framework for Quantile Residual Lifetime

Estimating prognosis conditional on surviving an initial high-risk period is crucial in clinical research. Yet, standard metrics such as hazard ratios are often difficult to interpret, while mean-based summaries are sensitive to outliers and censoring. We propose a formal causal framework for estimating quantiles of residual lifetime among individuals surviving to a landmark time $t_0$. Our primary estimand, the "Observed Survivor Quantile Contrast" (OSQC), targets pragmatic prognostic differences within the observed survivor population. To estimate the OSQC, we develop a doubly robust estimator that combines propensity scores, outcome regression, and inverse probability of censoring weights, ensuring consistency under confounding and informative censoring provided that the censoring model is correctly specified and at least one additional nuisance model is correctly specified. Recognizing that the OSQC conflates causal efficacy and compositional selection, we also introduce a reweighting-based supplementary estimator for the "Principal Survivor Quantile Contrast" (PSQC) to disentangle these mechanisms under stronger assumptions. Extensive simulations demonstrate the robustness of the proposed estimators and clarify the role of post-treatment selection. We illustrate the framework using data from the SUPPORT study to assess the impact of right heart catheterization on residual lifetime among intensive care unit survivors, and from the NSABP B-14 trial to examine post-surgical prognosis under adjuvant tamoxifen therapy across multiple landmark times.

stat.ME↗

StaRQR-K: False Discovery Rate Controlled Regional Quantile Regression

Quantifying how genomic features influence different parts of an outcome distribution requires statistical tools that go beyond mean regression, especially in ultrahigh-dimensional settings. Motivated by the study of LINE-1 activity in cancer, we propose StaRQR-K, a stabilized regional quantile regression framework with model-X knockoffs for false discovery rate control. StaRQR-K identifies CpG sites whose methylation levels are associated with specific quantile regions of an outcome, allowing detection of heterogeneous and tail-sensitive effects. The method combines an efficient regional quantile sure independence screening procedure with a winsorizing-based model-X knockoff filter, providing false discovery rate (FDR) control for regional quantile regression. Simulation studies show that StaRQR-K achieves valid FDR control and substantially higher power than existing approaches. In an application to The Cancer Genome Atlas head and neck cancer cohort, StaRQR-K reveals quantile-region-specific associations between CpG methylation and LINE-1 activity that improve out-of-sample prediction and highlight genomic regions with known functional relevance.

stat.ME↗

Decomposition of Longitudinal Disparities: an Application to the Fetal Growth-Singletons Study

Addressing health disparities among different demographic groups is a key challenge in public health. Despite many efforts, there is still a gap in understanding how these disparities unfold over time. Our paper focuses on this overlooked longitudinal aspect, which is crucial in both clinical and public health settings. In this paper, we introduce a longitudinal disparity decomposition method that decomposes disparities into three components: the explained disparity linked to differences in the exploratory variables' conditional distribution when the modifier distribution is identical between majority and minority groups, the explained disparity that emerges specifically from the unequal distribution of the modifier and its interaction with covariates, and the unexplained disparity. The proposed method offers a dynamic alternative to the traditional Peters-Belson decomposition approach, tackling both the potential reduction in disparity if the covariate distributions of minority groups matched those of the majority group and the evolving nature of disparity over time. We apply the proposed approach to a fetal growth study to gain insights into disparities between different race/ethnicity groups in fetal developmental progress throughout the course of pregnancy.

stat.AP↗