SearcharxivSearch

arXiv subjects

Taekwon Hong

Publications and source records attributed to Taekwon Hong.

5 recordsLinked to original sources

Beyond Point Estimates: Reliable Evaluation of Prediction Performance Metrics under Clustered Data

Prediction performance metrics such as accuracy and the F1 score are typically reported as single numbers, with no measure of uncertainty. The omission has been tolerable in exploratory settings, where model evaluation is used for informal comparison rather than formal decision-making. But as machine learning is deployed in real-world applications, evaluation results are increasingly used to support binary decisions -- whether a model meets a required standard or not -- making uncertainty quantification essential. The problem is compounded when data are dependent, as in repeated measurements, clustered subjects, or time series, where variability is harder to assess and easy to underestimate. We develop a unified framework that links a broad class of performance metrics through their representation as smooth functionals of confusion-matrix probabilities. This representation allows the use of the cluster-robust sandwich variance estimator to obtain asymptotically valid confidence intervals, hypothesis tests, and paired model comparisons for both binary and multiclass problems under clustered data. We also provide power and sample size approximations based on pilot data, enabling principled study design for model evaluation. Simulations show that the proposed methods achieve near-nominal coverage across a range of dependence structures, while naive methods underestimate variability. A real-data application further illustrates how accounting for clustering can materially change conclusions. These results offer a practical foundation for uncertainty quantification and study design in prediction performance evaluation, in settings where decisions should be justified under dependent and clustered data.

stat.ME

Bayesian Nonparametric Causal Inference for Quantile Residual Life: An Application to Alzheimer's Disease

In Alzheimer's disease research, a clinically important question is how much longer individuals would remain dementia-free beyond a given time under different baseline amyloid statuses. We address this question using observational data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) treating baseline amyloid status as the exposure. Estimation is challenging because amyloid status is confounded, time to dementia onset is heterogeneous and heavily right censored, and the target population depends on joint potential event times. At each time point, we consider the always-survivor principal stratum comprising individuals who would remain dementia-free under both amyloid status and estimate quantile contrasts in residual time to dementia onset. We model the joint distribution of event time, exposure, and baseline covariates using an enriched Dirichlet process mixture and conduct posterior inference via Bayesian g-computation. The framework accommodates partially observed covariates under a within-subcluster missing-at-random assumption, estimates contrasts across multiple time points and quantiles from one posterior fit and supports sensitivity analyses for unmeasured confounding, cross-world dependence, and informative censoring. Simulations show favorable finite-sample performance under heterogeneity and heavy censoring. In ADNI, residual time to dementia onset was shorter under elevated than non-elevated baseline amyloid status, both overall and within baseline subgroups.

stat.AP

Estimator-Aligned Prospective Sample Size Determination for Designs Using Inverse Probability of Treatment Weighting

In observational studies, accurately characterizing variance is critical for sample size determination, yet unaccounted-for variability from propensity score estimation and the resulting weights limit the accuracy of standard variance approximations for design. Existing approaches often rely on heuristics or randomized controlled trial (RCT) formulas that treat weights as fixed, potentially misaligning prospective design with the causal estimator used at analysis. We propose an estimator-aligned framework for prospective sample size determination based on generalized estimating equations (GEE) and stacked M-estimation. By merging the propensity score model and marginal structural model (MSM) into a single system of estimating equations, the method propagates nuisance-model uncertainty and directly targets the large-sample variance of the IPTW estimator. For study planning, we estimate a pilot-based large-sample variance factor and introduce a bootstrap stabilization procedure that accounts for both within- and between-pilot variability. The framework applies uniformly across binary, count, and continuous outcomes through link-specific GEE representations under a common design principle. Simulation studies motivated by post-marketing safety and healthcare cost applications demonstrate that anchoring design to this variance improves power calibration relative to conventional RCT-style formulas, particularly in settings with weight instability, outcome sparsity, or heavy-tailed variability.

stat.ME

A Causal Framework for Quantile Residual Lifetime

Estimating prognosis conditional on surviving an initial high-risk period is crucial in clinical research. Yet, standard metrics such as hazard ratios are often difficult to interpret, while mean-based summaries are sensitive to outliers and censoring. We propose a formal causal framework for estimating quantiles of residual lifetime among individuals surviving to a landmark time $t_0$. Our primary estimand, the "Observed Survivor Quantile Contrast" (OSQC), targets pragmatic prognostic differences within the observed survivor population. To estimate the OSQC, we develop a doubly robust estimator that combines propensity scores, outcome regression, and inverse probability of censoring weights, ensuring consistency under confounding and informative censoring provided that the censoring model is correctly specified and at least one additional nuisance model is correctly specified. Recognizing that the OSQC conflates causal efficacy and compositional selection, we also introduce a reweighting-based supplementary estimator for the "Principal Survivor Quantile Contrast" (PSQC) to disentangle these mechanisms under stronger assumptions. Extensive simulations demonstrate the robustness of the proposed estimators and clarify the role of post-treatment selection. We illustrate the framework using data from the SUPPORT study to assess the impact of right heart catheterization on residual lifetime among intensive care unit survivors, and from the NSABP B-14 trial to examine post-surgical prognosis under adjuvant tamoxifen therapy across multiple landmark times.

stat.ME

Multivariate Zero-Inflated Causal Model for Regional Mobility Restriction Effects on Consumer Spending

The COVID-19 pandemic presents challenges to both public health and the economy. Our objective is to examine how household expenditure, a significant component of private demand, reacts to changes in mobility. This investigation is crucial for developing policies that balance public health and the economic and social impacts. We utilize extensive scanner data from a major retail chain in India and Google mobility data to address this important question. However, there are a few challenges, including outcomes with excessive zeros and complicated correlations, time-varying confounding, and irregular observation times. We propose incorporating a multiplicative structural nested mean model with inverse intensity weighting techniques to tackle these challenges. Our framework allows semiparametric/nonparametric estimation for nuisance functions. The resulting rate doubly robust estimator enables the use of a conventional sandwich variance estimator without taking into account the variability introduced by these flexible estimation methods. We demonstrate the properties of our method theoretically and further validate it through simulation studies. Using the Indian consumer spending data and Google mobility data, our method reveals that the substantial reduction in mobility has a significant impact on consumers' fresh food expenditure.

stat.ME