Searcharxiv⌕ Search

arXiv subjects

Michael R. Kosorok

Publications and source records attributed to Michael R. Kosorok.

At least 19 recordsLinked to original sources

Latent Utility Q-Learning for Preference-Adaptive Dynamic Treatment Regimes

Optimizing individualized treatment sequences for patients who weigh multiple, competing outcomes differently poses a challenge for dynamic treatment regime (DTR) methods, which typically assume a single univariate outcome. We propose Latent Utility Q-Learning (LUQ-Learning), which estimates DTRs optimizing patient-specific preference-weighted combinations of multivariate outcomes $\mathbf{Y}\in\mathbb{R}^d$ across $K$ decision points. A conditional mean factorization decouples preference estimation from outcome regression, enabling flexible, modular learning under imperfectly observed and heterogeneous preferences without requiring explicit outcome ranking by patients. We establish consistency of the estimated value function and derive unified $ε$-optimality guarantees that bound policy value loss in terms of posterior preference uncertainty, yielding interpretable criteria for data-driven policy selection. Simulations calibrated to Sequential Multiple Assignment Randomized Trials (SMARTs) demonstrate that LUQ-Learning outperforms Q-learning with naive outcome aggregation, last-reported satisfaction optimization, and existing preference-based methods.

stat.ML↗

Backward Bayesian Outcome Weighted Learning

A central objective of precision medicine is learning optimal dynamic treatment regimes (DTRs) from data. Classification-based methods, like outcome weighted learning (OWL) for single-stage and backward OWL (BOWL) for multi-stage problems, leverage machine learning to directly learn optimal DTRs. However, these methods lack a natural way to quantify uncertainty in treatment decisions at the individual level. In this paper, we extend Bayesian OWL, a Bayesian reformulation of OWL, to the multi-stage setting. We call this method backward Bayesian outcome weighted learning (BBOWL). Like BOWL, our method directly learns an optimal DTR via backward induction, and unlike existing methods, our approach propagates uncertainty backward through the DTR learning process and provides uncertainty quantification of individualized treatment recommendations. We present a theoretical justification of BBOWL and verify its performance via a simulation study.

stat.ME↗

Bayesian Mediation Analysis for Individualized Treatment Rules

The value of an individualized treatment rule (ITR), defined as the expected outcome under treatment assignment according to the rule, is useful for assessing average clinical benefit but does not explain how the benefit of a rule is generated. We propose a causal mediation framework for decomposing the value contrast between a prespecified candidate ITR and a clinically meaningful reference rule into direct and indirect components. Using rule-specific nested potential outcomes, we define natural direct and indirect rule effects that quantify the extent to which the improvement in value arises through pathways operating directly on the outcome versus through a specified mediator. We give identification conditions under which these components are identified by a rule-level mediation g-formula. For estimation, we adapt Bayesian causal mediation forests to obtain posterior inference for the value contrast and its path-specific components. Our simulations demonstrate that the proposed estimator achieved near-nominal credible interval coverage with decreasing bias and root mean squared error as sample size increased in settings with varying direct and mediated contributions. We further illustrate the method using data from the TRIUMPH trial, decomposing the cognitive benefit of a lifestyle intervention rule through candidate neurovascular, cardiorespiratory, and behavioral mediators. The proposed framework complements optimal ITR learning with explanation using mediation, providing a natural approach for mechanistic evaluation of ITRs.

stat.ME↗

The Fidelity and Feedback Traps: The Case for Health Digital Twins as Modular Evolving Causal Systems

Digital twins for health may be used to compare treatments, project patient trajectories, and support clinical decisions. While related to mechanical digital twins, those initially developed for engineering applications, replicating the mechanical digital twin architecture and goals may fail in health for two reasons. The fidelity trap is the belief that an accurate model can answer what-if questions by virtue of its accuracy. Prediction and counterfactual reasoning are different tasks, and a twin that can fit past trajectories well may miss the mark when ranking treatments. The feedback trap arises when the twin updates on data its own recommendations helped generate. Refitting in this way can recover a biased relationship and grow more confident even as data grows thinner. We contend that health digital twins should be conceived as causally valid, modular, and evolving systems. Modularity isolates the data and models needed for interventional recommendations, causal validity supports such claims, and governed evolution updates the twin while accounting for how its recommendations reshape the data. We conclude that the standard for a health twin should be how well it supports decisions in the world it helps create, not how faithfully it reproduces the world it observes.

stat.OT↗

Distributional Random Forests for Complex Survey Designs

We study estimation of the conditional law $P(Y|X = x)$ and continuous measurable maps of it when $Y \in \mathcal{Y}$ takes values in a locally compact Polish space (e.g., $\mathbb{R}^d$), $X \in \mathbb{R}^p$, and the observations arise from a complex survey design: a single- or multi-stage sampling scheme that may involve unequal selection, stratification, and clustering. We propose a survey-calibrated distributional random forest (SDRF) that incorporates complex-design features via the pseudo-population bootstrap, PSU-level honesty, and a Maximum Mean Discrepancy (MMD) split criterion computed from kernel mean embeddings of design-weighted node distributions. We provide a framework for analyzing forest-based estimators under various survey designs; establish consistency for both finite- and super-population conditional laws under explicit conditions on the design, kernel, resampling multipliers, and tree partitions. As far as we are aware, these are the first results on model-free estimation of conditional distributions under survey designs. Simulations under a stratified two-stage cluster design expose the systematic bias incurred by ignoring survey structure. We illustrate the broad applicability of SDRF on NHANES, estimating the conditional joint tolerance regions for two diabetes biomarkers, revealing subgroup-level distributional heterogeneity relevant to diabetes risk profiling in the U.S. population.

stat.ME↗

Weighted NPMLE for the Marginal Mean of Recurrent Events with a Competing Terminal Event

Regression modeling of recurrent and terminal events continues to present methodological challenges in survival analysis. Existing approaches either make unverifiable assumptions about the dependency structure between the two event types or rely on the proportional intensity assumption for the marginal mean. A semiparametric regression model is proposed that is based on a novel weighted likelihood function, thereby targeting directly the marginal mean of the recurrent event. Our general model captures a large class of semiparametric regression models and accommodates external time-dependent covariate effects on the marginal mean intensity. We establish the consistency and asymptotic normality of the estimators and propose a sandwich estimator of the variance. We propose a novel simulation procedure that directly targets the marginal mean intensity of the recurrent events. In simulation studies, we demonstrate a strong performance of the weighted NPMLE under independent right-censoring. The practical utility of the proposed methodology is demonstrated through application to data from the STATCOPE trial, a large randomized clinical trial that investigated the efficacy of simvastatin for COPD exacerbations. We provide personalized predictions for the number of exacerbations and reassess the effect of simvastatin treatment, accounting for death as a competing terminal event for patients with GOLD stage 4.

stat.ME↗

A Review of Methods and Practices for Missing Data in Sequential Multiple Assignment Randomized Trials (SMARTs): An Ancillary Study of a Scoping Review

Background: Missing data poses an acute threat to sequential multiple assignment randomized trial (SMART) analyses because of the sequential treatment structure and response-dependent re-randomization. Objectives: This study aimed to (1) review the current statistical methods for handling missing data in SMARTs, and (2) characterize how missing data is reported and handled in published SMARTs. Methods: We conducted a narrative review of statistical methods developed for missing data in SMARTs. Additionally, we conducted a pre-specified secondary extraction of a previously published scoping review of SMARTs focused on missing data. Extraction captured attrition rates, methods for handling missingness, and planned versus performed missing data analyses. Results: Seven methodological papers were identified; nearly all assume missing at random (MAR), and only one addresses the full set of SMART-specific missingness types. Across 30 published SMARTs, median overall attrition was 18.1% (range 0.6%-56.5%). Methods used to address missing data were described in 80% of the manuscripts; mixed-model methods were most common (30%). Among 14 studies with paired protocols, sensitivity analyses were pre-specified in 2 (14%). Conclusions: SMART-specific methodology for missing data is limited, and a substantial gap exists between available methodology and current SMART practice.

stat.ME↗

Linear Regression Using Principal Components from General Hilbert-Space-Valued Covariates

We introduce Adaptive Subspace PCA (AS-PCA), a framework for principal component analysis of random elements in a general separable Hilbert space. AS-PCA projects the covariance operator onto a data-adaptive finite-dimensional subspace prior to eigendecomposition, requiring no kernel specification and accommodating multi-dimensional functional objects including images and surfaces. Under the second-moment condition, we prove a Donsker theorem for Hilbert-space-valued empirical processes and use it to establish uniform consistency and joint Gaussian limits for the leading eigenpairs. A data-driven diagnostic verifies projection accuracy, and a consistent proportion-of-variance-explained rule selects the number of components. Building on AS-PCA, we construct Hilbert-Space Principal Component Regression (HS-PCR) for models combining Euclidean and Hilbert-space-valued covariates. The HS-PCR estimator is root-$n$ consistent and asymptotically normal, with an explicit influence function decomposition accounting for eigenfunction estimation uncertainty. Both nonparametric and wild bootstrap procedures are shown to be asymptotically valid. Simulations with two- and three-dimensional imaging predictors confirm accurate eigenstructure recovery and nominal bootstrap coverage. HS-PCR is applied to Alzheimer's Disease Neuroimaging Initiative data in regression and precision-medicine settings.

math.ST↗

Provable Offline Reinforcement Learning for Structured Cyclic MDPs

We introduce a novel cyclic Markov decision process (MDP) framework for multi-step decision problems with heterogeneous stage-specific dynamics, transitions, and discount factors across the cycle. In this setting, offline learning is challenging: optimizing a policy at any stage shifts the state distributions of subsequent stages, propagating mismatch across the cycle. To address this, we propose a modular structural framework that decomposes the cyclic process into stage-wise sub-problems. While generally applicable, we instantiate this principle as CycleFQI, an extension of fitted Q-iteration enabling theoretical analysis and interpretation. It uses a vector of stage-specific Q-functions, tailored to each stage, to capture within-stage sequences and transitions between stages. This modular design enables partial control, allowing some stages to be optimized while others follow predefined policies. We establish finite-sample suboptimality error bounds and derive global convergence rates under Besov regularity, demonstrating that CycleFQI mitigates the curse of dimensionality compared to monolithic baselines. Additionally, we propose a sieve-based method for asymptotic inference of optimal policy values under a margin condition. Experiments on simulated and real-world Type 1 Diabetes data sets demonstrate CycleFQI's effectiveness.

stat.ML↗

Medical Knowledge Integration into Reinforcement Learning Algorithms for Dynamic Treatment Regimes

The goal of precision medicine is to provide individualized treatment at each stage of chronic diseases, a concept formalized by Dynamic Treatment Regimes (DTR). These regimes adapt treatment strategies based on decision rules learned from clinical data to enhance therapeutic effectiveness. Reinforcement Learning (RL) algorithms allow to determine these decision rules conditioned by individual patient data and their medical history. The integration of medical expertise into these models makes possible to increase confidence in treatment recommendations and facilitate the adoption of this approach by healthcare professionals and patients. In this work, we examine the mathematical foundations of RL, contextualize its application in the field of DTR, and present an overview of methods to improve its effectiveness by integrating medical expertise.

stat.ME↗

Constructing a T-test for Value Function Comparison of Individualized Treatment Regimes in the Presence of Multiple Imputation for Missing Data

Optimal individualized treatment decision-making has improved health outcomes in recent years. The value function is commonly used to evaluate the goodness of an individualized treatment decision rule. Despite recent advances, comparing value functions between different treatment decision rules or constructing confidence intervals around value functions remains difficult. We propose a t-test based method applied to a test set that generates valid p-values to compare value functions between a given pair of treatment decision rules when some of the data are missing. We demonstrate the ease in use of this method and evaluate its performance via simulation studies and apply it to the China Health and Nutrition Survey data.

stat.ME↗

Uncertainty quantification for intervals

Data following an interval structure are increasingly prevalent in many scientific applications. In medicine, clinical events are often monitored between two clinical visits, making the exact time of the event unknown and generating outcomes with a range format. As interest in automating healthcare decisions grows, uncertainty quantification via predictive regions becomes essential for developing reliable and trustworthy predictive algorithms. However, the statistical literature currently lacks a general methodology for interval targets, especially when these outcomes are incomplete due to censoring. We propose an uncertainty quantification algorithm for interval responses and establish its theoretical properties using empirical process arguments based on a newly developed class of functions specifically designed for these interval data structures. Although this paper primarily focuses on deriving predictive regions for interval-censored data, the approach can also be applied to other statistical modeling tasks, such as goodness-of-fit assessments. Finally, the applicability of the method is demonstrated through simulations, showing up to a 60\% improvement in conditional coverage. Our new algorithm is also applied to various biomedical contexts, including two clinical examples: i) sleep duration and its association with cardiovascular diseases, and ii) survival time in relation to physical activity levels.

stat.ME↗

Statistical Design and Rationale of the Biomarkers for Evaluating Spine Treatments (BEST) Trial

Chronic low back pain (cLBP) is a prevalent condition with profound impacts on functioning and quality of life. While multiple evidence-based treatments exist, they all have modest average treatment effects$\unicode{x2013}$potentially due to individual variation in treatment response and the diverse etiologies of cLBP. This multi-site sequential, multiple-assignment randomized trial (SMART) investigated four treatment modalities with two stages of randomization and aimed to enroll 630 protocol completers. The primary objective was to develop a precision medicine approach by estimating optimal treatment or treatment combinations based on patient characteristics and initial treatment response. The analysis strategy focuses on estimating interpretable dynamic treatment regimes and identifying subgroups most responsive to specific interventions. Broad eligibility criteria were implemented to enhance generalizability and recruitment, most notably that participants could be eligible to enroll even if they could not be assigned to one (but no more) of the study interventions. Enrolling participants with restrictions on the treatment they could be assigned necessitated modifications to standard minimization methods for balancing covariates. The BEST trial represents one of the largest SMARTs focused on clinical decision-making to date and the largest in cLBP. By collecting an extensive array of biomarker and phenotypic measures, this trial may identify potential treatment mechanisms and establish a more evidence-based approach to individualizing cLBP treatment in clinical practice.

stat.AP↗

An optimal dynamic treatment regime estimator for indefinite-horizon survival outcomes

We propose a new method in indefinite-horizon settings for estimating optimal dynamic treatment regimes for time-to-event outcomes. This method allows patients to have different numbers of treatment stages and is constructed using generalized survival random forests to maximize mean survival time. We use summarized history and data pooling, preventing data from growing in dimension as a patient's decision points increase. The algorithm operates through model re-fitting, resulting in a single model optimized for all patients and all stages. We derive theoretical properties of the estimator such as consistency of the estimator and value function and characterize the number of refitting iterations needed. We also conduct a simulation study of patients with a flexible number of treatment stages to examine finite-sample performance of the estimator. Finally, we illustrate use of the algorithm using administrative insurance claims data for pediatric Crohn's disease patients.

stat.ME↗

Using Statistical Precision Medicine to Identify Optimal Treatments in a Heart Failure Setting

Identifying optimal medical treatments to improve survival has long been a critical goal of pharmacoepidemiology. Traditionally, we use an average treatment effect measure to compare outcomes between treatment plans. However, new methods leveraging advantages of machine learning combined with the foundational tenets of causal inference are offering an alternative to the average treatment effect. Here, we use three unique, precision medicine algorithms (random forests, residual weighted learning, efficient augmentation relaxed learning) to identify optimal treatment rules where patients receive the optimal treatment as indicated by their clinical history. First, we present a simple hypothetical example and a real-world application among heart failure patients using Medicare claims data. We next demonstrate how the optimal treatment rule improves the absolute risk in a hypothetical, three-modifier setting. Finally, we identify an optimal treatment rule that optimizes the time to outcome in a real-world heart failure setting. In both examples, we compare the average time to death under the optimized, tailored treatment rule with the average time to death under a universal treatment rule to show the benefit of precision medicine methods. The improvement under the optimal treatment rule in the real-world setting is greatest (additional ~9 days under the tailored rule) for survival time free of heart failure readmission.

stat.AP↗

Optimal individualized treatment regimes for survival data with competing risks

Precision medicine leverages patient heterogeneity to estimate individualized treatment regimens, formalized, data-driven approaches designed to match patients with optimal treatments. In the presence of competing events, where multiple causes of failure can occur and one cause precludes others, it is crucial to assess the risk of the specific outcome of interest, such as one type of failure over another. This helps clinicians tailor interventions based on the factors driving that particular cause, leading to more precise treatment strategies. Currently, no precision medicine methods simultaneously account for both survival and competing risk endpoints. To address this gap, we develop a nonparametric individualized treatment regime estimator. Our two-phase method accounts for both overall survival from all events as well as the cumulative incidence of a main event of interest. Additionally, we introduce a multi-utility value function that incorporates both outcomes. We develop random survival and random cumulative incidence forests to construct individual survival and cumulative incidence curves. Simulation studies demonstrated that our proposed method performs well, which we applied to a cohort of peripheral artery disease patients at high risk for limb loss and mortality.

stat.ME↗

Effects Among the Affected

Many interventions are both beneficial to initiate and harmful to stop. Traditionally, to determine whether to deploy that intervention in a time-limited way depends on if, on average, the increase in the benefits of starting it outweigh the increase in the harms of stopping it. We propose a novel causal estimand that provides a more nuanced understanding of the effects of such treatments, particularly, how response to an earlier treatment (e.g., treatment initiation) modifies the effect of a later treatment (e.g., treatment discontinuation), thus learning if there are effects among the (un)affected. Specifically, we consider a marginal structural working model summarizing how the average effect of a later treatment varies as a function of the (estimated) conditional average effect of an earlier treatment. We allow for estimation of this conditional average treatment effect using machine learning, such that the causal estimand is a data-adaptive parameter. We show how a sequentially randomized design can be used to identify this causal estimand, and we describe a targeted maximum likelihood estimator for the resulting statistical estimand, with influence curve-based inference. Throughout, we use the Adaptive Strategies for Preventing and Treating Lapses of Retention in HIV Care trial (NCT02338739) as an illustrative example, showing that discontinuation of conditional cash transfers for HIV care adherence was most harmful among those who most had an increase in benefits from them initially.

stat.ME↗

Neural interval-censored survival regression with feature selection

Survival analysis is a fundamental area of focus in biomedical research, particularly in the context of personalized medicine. This prominence is due to the increasing prevalence of large and high-dimensional datasets, such as omics and medical image data. However, the literature on non-linear regression algorithms and variable selection techniques for interval-censoring is either limited or non-existent, particularly in the context of neural networks. Our objective is to introduce a novel predictive framework tailored for interval-censored regression tasks, rooted in Accelerated Failure Time (AFT) models. Our strategy comprises two key components: i) a variable selection phase leveraging recent advances on sparse neural network architectures, ii) a regression model targeting prediction of the interval-censored response. To assess the performance of our novel algorithm, we conducted a comprehensive evaluation through both numerical experiments and real-world applications that encompass scenarios related to diabetes and physical activity. Our results outperform traditional AFT algorithms, particularly in scenarios featuring non-linear relationships.

stat.ML↗