SearcharxivSearch

arXiv subjects

Yating Zou

Publications and source records attributed to Yating Zou.

6 recordsLinked to original sources

Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach

Missing data constitute a pervasive challenge in empirical research. Consequently, there is an ever-growing number of methods designed to address this challenge, with multiple imputation and inverse probability weighting the dominant strategies. Despite this, theoretical guarantees remain limited, particularly in the challenging case of non-monotone missing at random (MAR). When guarantees exist, they are often confined to simplified settings such as missing completely at random, monotone or block-wise missingness, or rest on restrictive assumptions about the missingness mechanism. In this paper, we utilize the theory of sieve maximum likelihood to establish a general rate of convergence under MAR that requires no modeling of the missingness mechanism and no restriction on the configuration of missing patterns, beyond MAR itself and a natural positivity condition. Applying this result to density estimation, we show that the complete-data density can be estimated at the minimax rate over a H\"older class, up to a logarithmic factor, for any prescribed smoothness level. The missingness does not affect the rate and enters only through a constant. The estimator is approximated in practice by a simple expectation-maximization (EM) algorithm operating on the incomplete data directly. In simulations, it performs comparably to the kernel density estimator supplied with the complete data across a wide range of missingness levels.

stat.ME

The Fidelity and Feedback Traps: The Case for Health Digital Twins as Modular Evolving Causal Systems

Digital twins for health may be used to compare treatments, project patient trajectories, and support clinical decisions. While related to mechanical digital twins, those initially developed for engineering applications, replicating the mechanical digital twin architecture and goals may fail in health for two reasons. The fidelity trap is the belief that an accurate model can answer what-if questions by virtue of its accuracy. Prediction and counterfactual reasoning are different tasks, and a twin that can fit past trajectories well may miss the mark when ranking treatments. The feedback trap arises when the twin updates on data its own recommendations helped generate. Refitting in this way can recover a biased relationship and grow more confident even as data grows thinner. We contend that health digital twins should be conceived as causally valid, modular, and evolving systems. Modularity isolates the data and models needed for interventional recommendations, causal validity supports such claims, and governed evolution updates the twin while accounting for how its recommendations reshape the data. We conclude that the standard for a health twin should be how well it supports decisions in the world it helps create, not how faithfully it reproduces the world it observes.

stat.OT

Distributional Random Forests for Complex Survey Designs

We study estimation of the conditional law $P(Y|X = x)$ and continuous measurable maps of it when $Y \in \mathcal{Y}$ takes values in a locally compact Polish space (e.g., $\mathbb{R}^d$), $X \in \mathbb{R}^p$, and the observations arise from a complex survey design: a single- or multi-stage sampling scheme that may involve unequal selection, stratification, and clustering. We propose a survey-calibrated distributional random forest (SDRF) that incorporates complex-design features via the pseudo-population bootstrap, PSU-level honesty, and a Maximum Mean Discrepancy (MMD) split criterion computed from kernel mean embeddings of design-weighted node distributions. We provide a framework for analyzing forest-based estimators under various survey designs; establish consistency for both finite- and super-population conditional laws under explicit conditions on the design, kernel, resampling multipliers, and tree partitions. As far as we are aware, these are the first results on model-free estimation of conditional distributions under survey designs. Simulations under a stratified two-stage cluster design expose the systematic bias incurred by ignoring survey structure. We illustrate the broad applicability of SDRF on NHANES, estimating the conditional joint tolerance regions for two diabetes biomarkers, revealing subgroup-level distributional heterogeneity relevant to diabetes risk profiling in the U.S. population.

stat.ME

Power Delivery for Cryogenic Scalable Quantum Applications: Challenges and Opportunities

Quantum technologies offer unprecedented capabilities in computation and secure information transfer. Their implementation requires qubits to operate at cryogenic temperatures (CT) while control and readout electronics typically still remains at room temperature (RT). As systems scale to millions of qubits, the electronics should also operate at CT to avoid a wiring bottleneck. However, wired power transfer from RT for such electronics introduces severe challenges, including thermal load between cooling stages, Joule heating, noise coupling, and wiring scalability. This paper addresses those challenges by evaluating several candidate architectures for scalable power transfer in the dilution frige: high-voltage (HV) wired power transfer, radiative wireless transfer, non-radiative wireless transfer, and a hybrid HV and non-radiative transfer. These architectures are analyzed in terms of thermal load, power loss, heating, coupling noise, power density, scalability, reliability, and complexity. Comparative analysis demonstrates the trade-offs among these architectures, while highlighting HV non-radiative transfer as a promising candidate for scalable quantum systems.

eess.SY

ROC Analysis with Covariate Adjustment Using Neural Network Models: Evaluating the Role of Age in the Physical Activity-Mortality Association

The receiver operating characteristic (ROC) curve and its summary measure, the Area Under the Curve (AUC), are well-established tools for evaluating the efficacy of biomarkers in biomedical studies. Compared to the traditional ROC curve, the covariate-adjusted ROC curve allows for individual evaluation of the biomarker. However, the use of machine learning models has rarely been explored in this context, despite their potential to develop more powerful and sophisticated approaches for biomarker evaluation. The goal of this paper is to propose a framework for neural network-based covariate-adjusted ROC modeling that allows flexible and nonlinear evaluation of the effectiveness of a biomarker to discriminate between two reference populations. The finite-sample performance of our method is investigated through extensive simulation tests under varying dependency structures between biomarkers, covariates, and referenced populations. The methodology is further illustrated in a clinically case study that assesses daily physical activity - measured as total activity time (TAC), a proxy for daily step count-as a biomarker to predict mortality at three, five and eight years. Analyzes stratified by sex and adjusted for age and BMI reveal distinct covariate effects on mortality outcomes. These results underscore the importance of covariate-adjusted modeling in biomarker evaluation and highlight TAC's potential as a functional capacity biomarker based on specific individual characteristics.

stat.ME

Latent Utility Q-Learning for Preference-Adaptive Dynamic Treatment Regimes

Optimizing individualized treatment sequences for patients who weigh multiple, competing outcomes differently poses a challenge for dynamic treatment regime (DTR) methods, which typically assume a single univariate outcome. We propose Latent Utility Q-Learning (LUQ-Learning), which estimates DTRs optimizing patient-specific preference-weighted combinations of multivariate outcomes $\mathbf{Y}\in\mathbb{R}^d$ across $K$ decision points. A conditional mean factorization decouples preference estimation from outcome regression, enabling flexible, modular learning under imperfectly observed and heterogeneous preferences without requiring explicit outcome ranking by patients. We establish consistency of the estimated value function and derive unified $\epsilon$-optimality guarantees that bound policy value loss in terms of posterior preference uncertainty, yielding interpretable criteria for data-driven policy selection. Simulations calibrated to Sequential Multiple Assignment Randomized Trials (SMARTs) demonstrate that LUQ-Learning outperforms Q-learning with naive outcome aggregation, last-reported satisfaction optimization, and existing preference-based methods.

stat.ML