SearcharxivSearch

arXiv subjects

Walter Dempsey

Publications and source records attributed to Walter Dempsey.

At least 19 recordsLinked to original sources

Estimation of Time-Varying Treatment Effects in a Joint Model for Longitudinal and Recurrent Event Outcomes in Mobile Health Data

Not only does mobile health technology enable researchers to track changes in multiple longitudinal outcomes of interest and to record the occurrence of health-related events over time, but it also allows for the delivery of repeated low-cost treatments directly to individuals in real time. We present a model-based approach for estimating the effect of repeatedly delivered treatments in a micro-randomized trial (MRT) via an extension of a joint longitudinal-survival model. We discuss different ways that these repeated treatment effects can be incorporated into the joint model; these different model specifications correspond to different mechanisms by which treatment is assumed to impact the longitudinal and event processes. Taking a Bayesian approach to inference, we model the association between repeated treatments, multiple longitudinally measured outcomes, and recurrent events. We also demonstrate how to calculate information criteria for model selection and present goodness-of-fit plots for assessing survival submodel calibration. We then illustrate the performance of our method via simulations and analysis of data collected in an MRT of substance use.

stat.ME

Data Integration for Estimating Subgroup-Specific Conditional Average Treatment Effects (CATEs) Using Coarsened External Information in Randomized Trials

Randomized controlled trials (RCTs) are often underpowered to detect treatment heterogeneity in subgroups defined by cross-classifications of multiple covariates, due to sparse sample sizes in some strata. External RCT data can help, but typically provide treatment effect estimates at a coarser level (e.g., by sex or race) rather than for the finer subgroups of interest (e.g., race-by-sex). We propose a novel James-Stein (JS)-type estimator that borrows strength from such coarsened external estimates to improve estimation of finer subgroup-specific conditional average treatment effects (CATEs) in an internal study, while accommodating potential incompatibility in marginal CATEs across populations. Based on asymptotic theory, we derive a practical analytic variance estimator for the JS estimator that exhibits acceptable empirical performance. Under mild conditions, we show that the proposed estimator uniformly dominates the ordinary least squares (OLS) estimator based on internal data regarding a weighted quadratic loss. Simulation studies demonstrate favorable performance compared with existing shrinkage methods, including empirical Bayes and generalized ridge estimators. We illustrate our method by estimating race-by-sex subgroup CATEs in a tirzepatide weight-loss trial (SURMOUNT-1), borrowing sex-specific and race-specific estimates from two previous semaglutide trials (STEP 1 and STEP 2). The proposed method detects a significantly larger treatment effect on percentage weight loss in the female-White subgroup than in the female-Asian subgroup, a difference not detected using internal data alone.

stat.ME

Evaluating time-varying treatment effects in hybrid SMART-MRT designs

Recently a new experimental approach, the hybrid experimental design (HED), was introduced to enable investigators to answer scientific questions about building behavioral interventions in which human-delivered and digital components are integrated and adapted on multiple timescales: slow (e.g., every few weeks) and fast (e.g., every few hours), respectively. An increasingly common HED involves the integration of the sequential, multiple assignment, randomized trial (SMART) with the micro-randomized trial (MRT), allowing investigators to answer scientific questions about potential synergistic effects of digital and human-delivered interventions. Approaches to formalize these questions in terms of causal estimands and associated data analytic methods are limited. In this paper, we formally define and assess these synergistic effects in hybrid SMART-MRTs on both proximal and distal outcomes. Practical utility is shown through the analysis of M-Bridge, a hybrid SMART-MRT aimed at reducing binge drinking among first-year college students.

stat.ME

Prediction Intervals for Individual Treatment Effects in a Multiple Decision Point Framework using Conformal Inference

Accurately quantifying uncertainty of individual treatment effects (ITEs) across multiple decision points is crucial for personalized decision-making in fields such as healthcare, finance, education, and online marketplaces. Previous work has focused on predicting non-causal longitudinal estimands or constructing prediction bands for ITEs using cross-sectional data based on exchangeability assumptions. We propose a novel method for constructing prediction intervals using conformal inference techniques for time-varying ITEs with weaker assumptions than prior literature. We guarantee a lower bound for coverage, which is dependent on the degree of non-exchangeability in the data. Although our method is broadly applicable across decision-making contexts, we support our theoretical claims with simulations emulating micro-randomized trials (MRTs) -- a sequential experimental design for mobile health (mHealth) studies. We demonstrate the practical utility of our method by applying it to a real-world MRT - the Intern Health Study (IHS).

stat.ME

Practical considerations when designing an online learning algorithm for an app-based mHealth intervention

The ubiquitous nature of mobile health (mHealth) technology has expanded opportunities for the integration of reinforcement learning into traditional clinical trial designs, allowing researchers to learn individualized treatment policies during the study. LowSalt4Life 2 (LS4L2) is a recent trial aimed at reducing sodium intake among hypertensive individuals through an app-based intervention. A reinforcement learning algorithm, which was deployed in one of the trial arms, was designed to send reminder notifications to promote app engagement in contexts where the notification would be effective, i.e., when a participant is likely to open the app in the next 30-minute and not when prior data suggested reduced effectiveness. Such an algorithm can improve app-based mHealth interventions by reducing participant burden and more effectively promoting behavior change. We encountered various challenges during the implementation of the learning algorithm, which we present as a template to solving challenges in future trials that deploy reinforcement learning algorithms. We provide template solutions based on LS4L2 for solving the key challenges of (i) defining a relevant reward, (ii) determining a meaningful timescale for optimization, (iii) specifying a robust statistical model that allows for automation, (iv) balancing model flexibility with computational cost, and (v) addressing missing values in gradually collected data.

stat.ME

Robust Bayesian Inference of Causal Effects via Randomization Distributions

We present a general framework for Bayesian inference of causal effects that delivers provably robust inferences founded on design-based randomization of treatments. The framework involves fixing the observed potential outcomes and forming a likelihood based on the randomization distribution of a statistic. The method requires specification of a treatment effect model; in many cases, however, it does not require specification of marginal outcome distributions, resulting in weaker assumptions compared to Bayesian superpopulation-based methods. We show that the framework is compatible with posterior model checking in the form of posterior-averaged randomization tests. We prove several theoretical properties for the method, including a Bernstein-von Mises theorem and large-sample properties of posterior expectations. In particular, we show that the posterior mean is asymptotically equivalent to Hodges-Lehmann estimators, which provides a bridge to many classical estimators in causal inference, including inverse-probability-weighted estimators and H\'ajek estimators. We evaluate the theory and utility of the framework in simulation and a case study involving a nutrition experiment. In the latter, our framework uncovers strong evidence of effect heterogeneity despite a lack of evidence for moderation effects. The basic framework allows numerous extensions, including the use of covariates, sensitivity analysis, estimation of assignment mechanisms, and generalization to nonbinary treatments.

stat.ME

Stable Central Limit Theorems for Discrete-Time Lag Martingale Difference Arrays: Applications to Dynamic Causal Inference

Recent work in dynamic causal inference introduced a class of discrete-time stochastic processes that generalize martingale difference sequences and arrays as follows: the random variates in each sequence have expectation zero given certain lagged filtrations but not given the natural filtration. We formalize this class of stochastic processes and prove stable central limit theorems (CLTs) via martingale-coboundary decomposition, leveraging the classical martingale CLT. We develop a variety of sufficient conditions, including conditions under which the limiting variance has a simple form that depends on variances and covariances of neighboring variates. We demonstrate the application of these results to inference for time-averaged treatment effects in switchback designs and present a simulation study supporting their validity. The CLTs enable various extensions to existing methodology for design-based approaches to dynamic causal inference, including time-lagged effects, random limiting variances, cross-unit dependence, and vector-valued estimands.

math.ST

Improving prediction in M-estimation by integrating external information from heterogeneous populations

A novel approach to improve prediction and inference in M-estimation by integrating external information from heterogeneous populations is proposed. Our method leverages joint asymptotics to combine estimates from external and internal datasets, where the external dataset provides auxiliary information about a subset of parameters of interest. We introduce a shrinkage estimator that combines internal and external estimates under a general class of transformations that ensure consistency across populations.

stat.ME

Evaluation of the HeartSteps Online Sampling Algorithm

Micro-randomized trials (MRTs), which sequentially randomize participants at multiple decision times, have gained prominence in digital intervention development. These sequential randomizations are often subject to certain constraints. In the MRT called HeartSteps V2V3, where an intervention is designed to interrupt sedentary behavior, two core design constraints need to be managed: an average of 1.5 interventions across days and the uniform delivery of interventions across decision times. Meeting both constraints, especially when the times allowed for randomization are not determined beforehand, is challenging. An online algorithm was implemented to meet these constraints in the HeartSteps V2V3 MRT. We present a case study using data from the HeartSteps V2V3 MRT, where we select appropriate metrics, discuss issues in making an accurate evaluation, and assess the algorithm's performance. Our evaluation shows that the algorithm performed well in meeting the two constraints. Furthermore, we identify areas for improvement and provide recommendations for designers of MRTs that need to satisfy these core design constraints.

stat.AP

Inference with Randomized Regression Trees

Regression trees are a popular machine learning algorithm that fit piecewise constant models by recursively partitioning the predictor space. This paper focuses on statistical inference for a data-dependent model obtained from a fitted regression tree. We introduce Randomized Regression Trees (RRT), a novel selective inference method that adds independent Gaussian noise to the gain function underlying the splitting rules of classic regression trees. The RRT method offers several advantages over existing methods. First, added randomization is used to obtain a closed-form pivot while accounting for the data-dependent tree structure. Second, RRT with a small amount of randomization achieves predictive accuracy similar to a model trained on the entire dataset, while also providing significantly more powerful inference than existing selective inference methods, such as data splitting. Third, RRT yields intervals that automatically adapt to the signal strength in the data. Our empirical analyses highlight these advantages of the RRT method and its ability to convert a purely predictive algorithm into a method capable of performing powerful inference in the non-linear tree model.

stat.ME

Selective Inference for Time-Varying Moderated Effects

Causal effect moderation investigates how the effect of interventions (or treatments) on outcome variables changes based on observed characteristics of individuals, known as potential effect moderators. With advances in data collection, datasets containing many observed features as potential moderators have become increasingly common. High-dimensional analyses often lack interpretability, with important moderators masked by noise, while low-dimensional, marginal analyses yield many false positives due to strong correlations with true moderators. In this paper, we propose a two-step method for selective inference on time-varying causal effect moderation that addresses the limitations of both high-dimensional and marginal analyses. Our method first selects a relatively smaller, more interpretable model to estimate a linear causal effect moderation using a Gaussian randomization approach. We then condition on the selection event to construct a pivot, enabling uniformly asymptotic semi-parametric inference in the selected model. Our numerical results show that our method achieves valid coverage rates, even when existing conditional methods and common sample splitting techniques fail. Moreover, our method yields shorter, bounded intervals, unlike existing methods that may produce infinitely long intervals.

stat.ME

Data integration methods for micro-randomized trials

Existing statistical methods for the analysis of micro-randomized trials (MRTs) are designed to estimate causal excursion effects using data from a single MRT. In practice, however, researchers can often find previous MRTs that employ similar interventions. In this paper, we develop data integration methods that capitalize on this additional information, leading to statistical efficiency gains. To further increase efficiency, we demonstrate how to combine these approaches according to a generalization of multivariate precision weighting that allows for correlation between estimates, and we show that the resulting meta-estimator possesses an asymptotic optimality property. We illustrate our methods in simulation and in a case study involving two MRTs in the area of smoking cessation.

stat.ME

Non-Stationary Latent Auto-Regressive Bandits

For the non-stationary multi-armed bandit (MAB) problem, many existing methods allow a general mechanism for the non-stationarity, but rely on a budget for the non-stationarity that is sub-linear to the total number of time steps $T$. In many real-world settings, however, the mechanism for the non-stationarity can be modeled, but there is no budget for the non-stationarity. We instead consider the non-stationary bandit problem where the reward means change due to a latent, auto-regressive (AR) state. We develop Latent AR LinUCB (LARL), an online linear contextual bandit algorithm that does not rely on the non-stationary budget, but instead forms good predictions of reward means by implicitly predicting the latent state. The key idea is to reduce the problem to a linear dynamical system which can be solved as a linear contextual bandit. In fact, LARL approximates a steady-state Kalman filter and efficiently learns system parameters online. We provide an interpretable regret bound for LARL with respect to the level of non-stationarity in the environment. LARL achieves sub-linear regret in this setting if the noise variance of the latent state process is sufficiently small with respect to $T$. Empirically, LARL outperforms various baseline methods in this non-stationary bandit problem.

cs.LG

Selective Inference for Sparse Graphs via Neighborhood Selection

Neighborhood selection is a widely used method used for estimating the support set of sparse precision matrices, which helps determine the conditional dependence structure in undirected graphical models. However, reporting only point estimates for the estimated graph can result in poor replicability without accompanying uncertainty estimates. In fields such as psychology, where the lack of replicability is a major concern, there is a growing need for methods that can address this issue. In this paper, we focus on the Gaussian graphical model. We introduce a selective inference method to attach uncertainty estimates to the selected (nonzero) entries of the precision matrix and decide which of the estimated edges must be included in the graph. Our method provides an exact adjustment for the selection of edges, which when multiplied with the Wishart density of the random matrix, results in valid selective inferences. Through the use of externally added randomization variables, our adjustment is easy to compute, requiring us to calculate the probability of a selection event, that is equivalent to a few sign constraints and that decouples across the nodewise regressions. Through simulations and an application to a mobile health trial designed to study mental health, we demonstrate that our selective inference method results in higher power and improved estimation accuracy.

stat.ME

Incorporating Auxiliary Variables to Improve the Efficiency of Time-Varying Treatment Effect Estimation

Contextual sensing and delivery of digital interventions to improve health outcomes have gained significant traction in behavioral and psychiatric studies. Micro-randomized trials (MRTs) are a common experimental design for obtaining data-driven evidence on the effectiveness of digital interventions where each individual is repeatedly randomized to receive treatments over numerous time points. Throughout the study, individual characteristics and contextual factors around randomization are collected, with some prespecified as moderators for assessing time-varying causal effect moderation. However, many additional measurements beyond these moderators often go underutilized. Some of these may influence treatment randomization or known to strongly moderate the treatment effect. Incorporating such auxiliary information into the estimation procedure can reduce chance imbalances and improve asymptotic estimation efficiency. In this work, we propose a method to adjust for auxiliary variables in consistently estimating time-varying intervention effects. The approach can also be extended to include post-treatment auxiliary variables when evaluating lagged treatment effects. Under specific conditions, local efficiency gains are guaranteed. We demonstrate the method's utility through simulation studies and an analysis of data from the Intern Health Study (NeCamp et al., 2020).

stat.ME

A Meta-Learning Method for Estimation of Causal Excursion Effects to Assess Time-Varying Moderation

Advances in wearable technologies and health interventions delivered by smartphones have greatly increased the accessibility of mobile health (mHealth) interventions. Micro-randomized trials (MRTs) are designed to assess the effectiveness of the mHealth intervention and introduce a novel class of causal estimands called "causal excursion effects." These estimands enable the evaluation of how intervention effects change over time and are influenced by individual characteristics or context. Existing methods for analyzing causal excursion effects assume known randomization probabilities, complete observations, and a linear nuisance function with prespecified features of the high dimensional observed history. However, in complex mobile systems, these assumptions often fall short: randomization probabilities can be uncertain, observations may be incomplete, and the granularity of mHealth data makes linear modeling difficult. To address this issue, we propose a flexible and doubly robust inferential procedure, called "DR-WCLS," for estimating causal excursion effects from a meta-learner perspective. We present the bidirectional asymptotic properties of the proposed estimators and compare them with existing methods both theoretically and through extensive simulations. The results show a consistent and more efficient estimate, even with missing observations or uncertain treatment randomization probabilities. Finally, the practical utility of the proposed methods is demonstrated by analyzing data from a multiinstitution cohort of first-year medical residents in the United States (NeCamp et al., 2020).

stat.ME

Exploring the big data paradox for various estimands using vaccination data from the global COVID-19 Trends and Impact Survey (CTIS)

Selection bias poses a challenge to statistical inference validity in non-probability surveys. This study compared estimates of the first-dose COVID-19 vaccination rates among Indian adults in 2021 from a large non-probability survey, COVID-19 Trends and Impact Survey (CTIS), and a small probability survey, the Center for Voting Options and Trends in Election Research (CVoter), against benchmark data from the COVID Vaccine Intelligence Network (CoWIN). Notably, CTIS exhibits a larger estimation error (0.39) compared to CVoter (0.16). Additionally, we investigated the estimation accuracy of the CTIS when using a relative scale and found a significant increase in the effective sample size by altering the estimand from the overall vaccination rate. These results suggest that the big data paradox can manifest in countries beyond the US and it may not apply to every estimand of interest.

stat.AP

Design of Experiments with Sequential Randomizations on Multiple Timescales: The Hybrid Experimental Design

Psychological interventions, especially those leveraging mobile and wireless technologies, often include multiple components that are delivered and adapted on multiple timescales (e.g., coaching sessions adapted monthly based on clinical progress, combined with motivational messages from a mobile device adapted daily based on the person's daily emotional state). The hybrid experimental design (HED) is a new experimental approach that enables researchers to answer scientific questions about the construction of psychological interventions in which components are delivered and adapted on different timescales. These designs involve sequential randomizations of study participants to intervention components, each at an appropriate timescale (e.g., monthly randomization to different intensities of coaching sessions and daily randomization to different forms of motivational messages). The goal of the current manuscript is twofold. The first is to highlight the flexibility of the HED by conceptualizing this experimental approach as a special form of a factorial design in which different factors are introduced at multiple timescales. We also discuss how the structure of the HED can vary depending on the scientific question(s) motivating the study. The second goal is to explain how data from various types of HEDs can be analyzed to answer a variety of scientific questions about the development of multi-component psychological interventions. For illustration we use a completed HED to inform the development of a technology-based weight loss intervention that integrates components that are delivered and adapted on multiple timescales.

stat.ME