SearcharxivSearch

arXiv subjects

Judith J. Lok

Publications and source records attributed to Judith J. Lok.

At least 19 recordsLinked to original sources

Regression Not-to-the-Mean: An Oddity of Regression, Illustrated with the Risk of Overdose Deaths

Recent works in econometrics have shown that there can be issues with applying a constant treatment effect model in longitudinal settings with staggered treatment and heterogeneous treatment effects. We focus on the issue that the estimated constant treatment effect may be a weighted average, with some negative weights, of treatment effects that are heterogeneous across treatment durations. When this issue arises, the estimated constant treatment effect and estimated heterogeneous treatment effects may result in conflicting results. Through the example of estimating the effect of drug-induced homicide (DIH) prosecutions reported by media on unintentional drug-overdose deaths in the United States, we illustrate how the negative weighting issue can lead to conflicting results in practice. Moreover, although research has shown that the negative weight issue may arise in linear regression models, we show this issue may also arise in logistic regression models. Using a linear link, we estimated a constant treatment effect risk ratio of 0.977 (95% CI:(0.866, 1.101)) and an average risk ratio of 0.728 (range: 0.507-0.979) over different treatment durations. Using a logistic link, we estimated a constant treatment risk ratio effect of 1.064 (95% CI: (0.972, 1.165)) and an average risk ratio of 0.739 (range: 0.538-1.008) over different treatment durations. Under both models, the estimated constant treatment effect is either smaller in magnitude or has a different sign than almost all estimated heterogeneous treatment effects, suggesting a negative weighting issue is present. Our results suggest additional care is needed when applying constant treatment effect models in longitudinal settings.

stat.AP

Addressing Confounding by Indication Through (Un)Measured Centre Characteristics in Learn-As-you-GO(LAGO) Trials

The Learn-As-you-Go (LAGO) design is an adaptive clinical trial design allowing modifications to multicomponent intervention packages across stages. Centers participate in more than one stage, as is common in large-scale implementation trials. In LAGO trials, center characteristics may act as confounders, predicting both the intervention package and the outcomes. We extend LAGO theory by introducing fixed center effects to control for confounding by indication through measured and unmeasured center characteristics. Conditioning on center characteristics by including fixed center effects ensures asymptotic results hold without requiring explicit characterization of unmeasured confounders. Our methods apply even with small numbers of centers. LAGO theory is established for continuous outcomes following a generalized linear model and binary outcomes following a logistic regression model, unifying theory across outcome types. Point- and interval estimators are derived, and consistency and asymptotic normality are established. Valid hypothesis tests for the overall intervention effect are provided, and the optimal intervention package minimizing cost subject to a target outcome mean is obtained via constrained optimization.

stat.ME

Optimizing Complex Health Intervention Packages through the Learn-As-you-GO (LAGO) Design

In the face of vast numbers of preventable deaths worldwide and gaping disparities in their distribution, we cannot afford to conduct null and inconclusive effectiveness and implementation trials of evidence-based interventions. The gold standard in biomedical research, the individually randomized clinical trial, is ill-suited as the primary tool for knowledge generation for contextually relevant, scalable, complex public health interventions of multi-component strategies. In this paper, we discuss the new Learn-As-you-GO (LAGO) design. In LAGO trials, the components of a complex intervention package are repeatedly optimized in pre-planned stages, until the package achieves its outcome and power goals with minimized cost and/or other optimization criteria, such as maximizing patient satisfaction. In this paper, the inputs to, and outputs of, LAGO are described, along with its general methodology. The methods are illustrated in the BetterBirth study, a large trial that aimed to reduce maternal and neonatal mortality in Uttar Pradesh, India, using the WHO essential birth practices checklist. Despite its scale, the BetterBirth study failed to demonstrate a significant effect of the intervention package on the primary health endpoint that included maternal mortality. We show how this unfortunate outcome could have been remedied had LAGO been used. LAGO is further illustrated through the discussion of several ongoing LAGO-informed implementation trials of HIV and non-communicable diseases in the United States and Sub-Saharan Africa. The Learn-As-you-GO (LAGO) design optimizes a complex, multi-level intervention for minimum cost, pre-specified power, and a pre-specified effectiveness goal, by adapting the intervention as the study is conducted, reducing risk of trial failure.

stat.ME

Learn-As-you-GO (LAGO) Trials: Optimizing Trials for Effectiveness and Power to Prevent Failed Trials

The Learn-As-you-GO (LAGO) design provides a rigorous framework for adapting the intervention package based on accumulating data while the trial is ongoing. This article improves the flexibility of the LAGO design by incorporating statistical power as an optimization criterion (power goal) in LAGO optimizations. We propose the unconditional and conditional power approaches to add a power goal. Both approaches estimate the power at the end of the LAGO trial using data from prior stages, and increase the power at the end of the LAGO trial when the original trial was underpowered. Including a power goal maintains the asymptotic properties of the estimators of the treatment effect while preserving the asymptotic level of the statistical test at the end of the trial. We illustrate the benefits of our methods through a retrospective application to the BetterBirth Study, a large-scale study of maternal-newborn care that failed to show a significant effect on its primary outcome. This analysis demonstrates how our methods could have led to more intensive interventions and potentially significant results. The LAGO design with power goal optimizations provides investigators with a powerful tool to reduce the risk of failed trials due to insufficient power.

stat.ME

Sex at birth could well be a biological coin toss.... Beware of conditioning on post-baseline information

Wang et al. (2025) use statistics to argue that sex at birth is not a biological coin toss, by noticing that repeated patterns such as Male Male Male and Female Female Female occur in the Nurses Health Study more often than patterns like Male Female Male, Male Female Female, Female Male Female, or Female Male Male. This letter shows that this over-representation is likely due to a statistical artifact, arising from parent preferences for mixed-sex children. As noticed in Angrist and Evans (1998) and supported by the data in Wang et al. (2025), parents are more likely to have a third child if their first two children are of the same sex. We show mathematically and statistically that mixed-sex preferences lead to the over-representation of patterns like Male Male Male and Female Female Female. In fact, the patterns seen in the Nurses Health Study are perfectly consistent with sex at birth being a random coin toss.

stat.AP

Causal indirect effect of an HIV curative treatment: mediators subject to an assay limit and measurement error

Causal mediation analysis decomposes the total effect of a treatment on an outcome into the indirect effect, operating through the mediator, and the direct effect, operating through other pathways. One can estimate only the pure indirect effect/indirect effect relative to no treatment, rather than the total effect by combining a hypothesized treatment effect on the mediator with outcome data without treatment. Furthermore, the mediation formula holds for the pure indirect effect (or the organic indirect effect relative to no treatment) regardless of whether there is an interaction between the treatment and mediator in the outcome model. This methodology holds significant promise in selecting prospective treatments based on their indirect effect for further evaluation in randomized clinical trials. We apply this methodology to assess which of two measures of HIV persistence is a more promising target for future HIV curative treatments. We combine a hypothesized treatment effect on two mediators, and outcome data without treatment, to compare the indirect effect of treatments targeting these mediators. Some HIV persistence measurements fall below the assay limit, leading to left-censored mediators. We address this by assuming the outcome model extends to mediators below the assay limit and use maximum likelihood estimation. To address measurement error in the mediators, we adjust our estimates. Using data from completed ACTG studies, we estimate the pure or organic indirect effect of potential curative HIV treatments on viral suppression through weeks 4 and 8 after HIV medication interruption, mediated by HIV persistence measures.

stat.AP

Estimating treatment effects from observational data under truncation by death using survival-incorporated quantiles

The issue of "truncation by death" commonly arises in clinical research: subjects may die before their follow-up assessment, resulting in undefined clinical outcomes. To address this issue, we focus on survival-incorporated quantiles -- quantiles of a composite outcome combining death and clinical outcomes -- to summarize the effect of treatment. Using inverse probability of treatment weighting (IPTW), we propose an estimator for survival-incorporated quantiles from observational data, applicable to settings of both point treatment and time-varying treatments. We establish consistency and asymptotic normality of the estimator under both the true and estimated propensity scores. While the variance properties of IPTW estimators for the mean have been studied, to our knowledge, this article is the first to show that the IPTW quantile estimator using the estimated propensity score yields lower asymptotic variance than the IPTW quantile estimator using the true propensity score. Extensive simulations show that survival-incorporated quantiles provide a simple and useful summary measure and confirm that using the estimated propensity score reduces the root mean square error. We apply our method to estimate the effect of statins on the change in cognitive function, incorporating death, using data from the Long Life Family Study (LLFS) -- a multicenter observational study of 4953 older adults with familial longevity. Our results indicate no significant difference in cognitive decline between statin users and non-users with a similar age- and sex-distribution at baseline. This study not only contributes to understand the cognitive effects of statins but also provides insights into analyzing clinical outcomes in the presence of death.

stat.ME

Demystified: double robustness with nuisance parameters estimated at rate n-to-the-1/4

Have you also been wondering what is this thing with double robustness and nuisance parameters estimated at rate n^(1/4)? It turns out that to understand this phenomenon one just needs the Middle Value Theorem (or a Taylor expansion) and some smoothness conditions. This note explains why under some fairly simple conditions, as long as the nuisance parameter theta in R^k is estimated at rate n^(1/4) or faster, 1. the resulting variance of the estimator of the parameter of interest psi in R^d does not depend on how the nuisance parameter theta is estimated, and 2. the sandwich estimator of the variance of psi-hat ignoring estimation of theta is consistent.

math.ST

The survival-incorporated median versus the median in the survivors or in the always-survivors: What are we measuring? And why?

Many clinical studies evaluate the benefit of a treatment based on both survival and other continuous/ordinal clinical outcomes, such as Quality of Life scores. In these studies, when subjects die before the follow-up assessment, the clinical outcomes become undefined and are truncated by death. Treating outcomes as "missing" or "censored" due to death can be misleading for treatment effect evaluation. We show that if we use the median in the survivors or in the always-survivors as estimands to summarize clinical outcomes, we may conclude that a trade-off exists between the probability of survival and good clinical outcomes, even in settings where both the probability of survival and the probability of any good clinical outcome are better for one treatment. Therefore, we advocate not always treating death as a mechanism through which clinical outcomes are missing, but rather as part of the outcome measure. To account for the survival status, we describe the survival-incorporated median as an alternative summary measure for outcomes in the presence of death. The survival-incorporated median is the threshold such that 50% of the population is alive with an outcome above that threshold. Through conceptual examples and an application to a prostate cancer treatment study, we show that the survival-incorporated median provides a simple and useful summary measure to inform clinical practice.

stat.AP

Learn-As-you-GO (LAGO) Trials: Optimizing Treatments and Preventing Trial Failure Through Ongoing Learning

It is well known that changing the intervention package while a trial is ongoing does not lead to valid inference using standard statistical methods. However, it is often necessary to adapt, tailor, or tweak a complex intervention package in public health implementation trials, especially when the intervention package does not have the desired effect. This article presents conditions under which the resulting analyses remain valid even when the intervention package is adapted while a trial is ongoing. Our results on such Learn-As-you-GO (LAGO) studies extend the theory of LAGO for binary outcomes following a logistic regression model (Nevo, Lok and Spiegelman, 2021) to LAGO for continuous outcomes under flexible conditional mean model. We derive point and interval estimators of the intervention effects and ensure the validity of hypothesis tests for an overall intervention effect. We develop a confidence set for the optimal intervention package, which achieves a pre-specified mean outcome while minimizing cost, and confidence bands for the mean outcome under all intervention package compositions. This work will be useful for the design and analysis of large-scale intervention trials where the intervention package is adapted, tailored, or tweaked while the trial is ongoing.

stat.ME

How estimating nuisance parameters can reduce the variance (with consistent variance estimation)

We often estimate a parameter of interest psi when the identifying conditions involve a nuisance parameter theta. Examples from causal inference are Inverse Probability Weighting, Marginal Structural Models and Structural Nested Models. To estimate treatment effects from observational data, these methods posit a (pooled) logistic regression model for the treatment and/or censoring probabilities and estimate these first. These methods are all based on unbiased estimating equations. First, we provide a general formula for the variance of the parameter of interest psi when the nuisance parameter theta is estimated in a first step, using the Partition Inverse Formula. Then, we present 4 results for estimators psi-hat based on unbiased estimating equations including a nuisance parameter theta which is estimated by solving (partial) score equations, if psi does not depend on theta. This regularly happens in causal inference if theta describes the treatment probabilities, in settings with missing data where theta describes the missingness probabilities, and settings with measurement error where theta describes the measurement error distribution. 1. Counter-intuitively, the limiting variance of psi-hat is typically smaller when theta is estimated, compared to if a known theta were plugged in. 2. If estimating theta is ignored, the resulting sandwich estimator for the variance of psi-hat is conservative. 3. A consistent estimator for the variance of psi-hat can provide results fast: no bootstrap. 4. If psi-hat with the true theta plugged in is efficient, the limiting variance of psi-hat does not depend on whether theta is estimated. To illustrate we use observational data to estimate 1. the effect of cazavi versus colistin in patients with resistant bacterial infections and 2. how the effect of one year of antiretroviral treatment depends on its initiation time in HIV-infected patients.

math.ST

Causal mediation analysis with mediator values below an assay limit

Causal indirect and direct effects provide an interpretable method for decomposing the total effect of an exposure on an outcome into the effect through a mediator and the effect through all other pathways. When the mediator is a biomarker, values can be subject to an assay lower limit. The mediator is affected by the treatment and is a putative cause of the outcome, so the assay lower limit presents a compounded problem in mediation analysis. We propose three approaches to estimate indirect and direct effects with a mediator subject to an assay limit: 1. extrapolation 2. numerical optimization and integration of the observed likelihood and 3. the Monte Carlo Expectation Maximization (MCEM) algorithm. Since the described methods solely rely on the so-called Mediation Formula, they apply to most approaches to causal mediation analysis: natural, separable, and organic indirect and direct effects. A simulation study compares the estimation approaches to imputing with half the assay limit. Using HIV interruption study data from the AIDS Clinical Trials Group described in [Li et al. 2016, AIDS; Lok \& Bosch 2021, Epidemiology], we illustrate our methods by estimating the organic/pure indirect effect of a hypothetical HIV curative treatment on viral suppression mediated by two HIV persistence measures: cell-associated HIV-RNA (N = 124) and single copy plasma HIV-RNA (N = 96).

stat.ME

Optimal estimation of coarse structural nested mean models with application to initiating ART in HIV infected patients

Coarse structural nested mean models are used to estimate treatment effects from longitudinal observational data. Coarse structural nested mean models lead to a large class of estimators. It turns out that estimates and standard errors may differ considerably within this class. We prove that, under additional assumptions, there exists an explicit solution for the optimal estimator within the class of coarse structural nested mean models. Moreover, we show that even if the additional assumptions do not hold, this optimal estimator is doubly-robust: it is consistent and asymptotically normal not only if the model for treatment initiation is correct, but also if a certain outcome-regression model is correct. We compare the optimal estimator to some naive choices within the class of coarse structural nested mean models in a simulation study. Furthermore, we apply the optimal and naive estimators to study how the CD4 count increase due to one year of antiretroviral treatment (ART) depends on the time between HIV infection and ART initiation in recently infected HIV infected patients. Both in the simulation study and in the application, the use of optimal estimators leads to substantial increases in precision.

math.ST

Causal organic indirect and direct effects: closer to Baron and Kenny, with a product method for binary mediators

Mediation analysis, which started with Baron and Kenny (1986), is used extensively by applied researchers. Indirect and direct effects are the part of a treatment effect that is mediated by a covariate and the part that is not. Subsequent work on natural indirect and direct effects provides a formal causal interpretation, based on cross-worlds counterfactuals: outcomes under treatment with the mediator set to its value without treatment. Organic indirect and direct effects (Lok 2016) avoid cross-worlds counterfactuals, using so-called organic interventions on the mediator while keeping the initial treatment fixed at treatment. Organic indirect and direct effects apply also to settings where the mediator cannot be set. In linear models where the outcome model does not have treatment-mediator interaction, both organic and natural indirect and direct effects lead to the same estimators as in Baron and Kenny (1986). Here, we generalize organic interventions on the mediator to include interventions combined with the initial treatment fixed at no treatment. We show that the product method holds in linear models for organic indirect and direct effects relative to no treatment even if there is treatment-mediator interaction. Moreover, we find a product method for binary mediators. Furthermore, we argue that the organic indirect effect relative to no treatment is very relevant for drug development. We illustrate the benefits of our approach by estimating the organic indirect effect of curative HIV-treatments mediated by two HIV-persistence measures, using ART-interruption data without curative HIV-treatments combined with an estimated/hypothesized effect of the curative HIV-treatments on these mediators.

stat.ME

Analysis of "Learn-As-You-Go" (LAGO) Studies

In learn-as-you-go (LAGO) adaptive studies, the intervention is a complex package consisting of multiple components, and is adapted in stages during the study based on past outcome data. This design formalizes standard practice, and desires for practice, in public health intervention studies. An effective intervention package is sought, while minimizing intervention package cost. When analyzing data from a learn-as-you-go study, the interventions in later stages depend upon the outcomes in the previous stages, violating standard statistical theory. We develop methods for estimating the intervention effects in a LAGO study. We prove consistency and asymptotic normality using a novel coupling argument, ensuring the validity of the test for the hypothesis of no overall intervention effect. We develop a confidence set for the optimal intervention package and confidence bands for the success probabilities under alternative package compositions. We illustrate our methods in the BetterBirth Study, which aimed to improve maternal and neonatal outcomes among 157,689 births in Uttar Pradesh, India through a complex, multi-component intervention package.

stat.ME

Doubly Robust Goodness-of-Fit Test of Coarse Structural Nested Mean Models with Application to Initiating combination antiretroviral treatment in HIV-Positive Patients

Coarse Structural Nested Mean Models (SNMMs) provide useful tools to estimate treatment effects from longitudinal observational data with time-dependent confounders. Coarse SNMMs lead to a large class of estimators,within which an optimal estimator can be derived under the conditions of well-specified models for the treatment effect, for treatment initiation, and for nuisance regression outcomes (Lok & Griner, 2015). The key assumption lies in a well-specified model for the treatment effect; however, there is no existing guidance to specify the treatment effect model, and model misspecification leads to biased estimators, preventing valid inference. To test whether the treatment effect model matches the data well, we derive a goodness-of-fit (GOF) test procedure based on overidentification restrictions tests (Sargan, 1958; Hansen, 1982). We show that our GOF statistic is doubly-robust in the sense that with a correct treatment effect model, if either the treatment initiation model or the nuisance regression outcome model is correctly specified, the GOF statistic has a Chi-squared limiting distribution with degrees of freedom equal to the number of overidentification restrictions. We demonstrate the empirical relevance of our methods using simulation designs based on an actual dataset. In addition, we apply the GOF test procedure to study how the initiation time of highly active antiretroviral treatment (HAART) after infection predicts the one-year treatment effect in HIV-positive patients with acute and early infection.

stat.ME

Defining and estimating causal direct and indirect effects when setting the mediator to specific values is not feasible

Natural direct and indirect effects decompose the effect of a treatment into the part that is mediated by a covariate (the mediator) and the part that is not. Their definitions rely on the concept of outcomes under treatment with the mediator "set" to its value without treatment. Typically, the mechanism through which the mediator is set to this value is left unspecified, and in many applications it may be challenging to fix the mediator to particular values for each unit or individual. Moreover, how one sets the mediator may affect the distribution of the outcome. This article introduces "organic" direct and indirect effects, which can be defined and estimated without relying on setting the mediator to specific values. Organic direct and indirect effects can be applied for example to estimate how much of the effect of some treatments for HIV/AIDS on mother-to-child transmission of HIV-infection is mediated by the effect of the treatment on the HIV viral load in the blood of the mother.

stat.ME

Mimicking counterfactual outcomes to estimate causal effects

In observational studies, treatment may be adapted to covariates at several times without a fixed protocol, in continuous time. Treatment influences covariates, which influence treatment, which influences covariates, and so on. Then even time-dependent Cox-models cannot be used to estimate the net treatment effect. Structural nested models have been applied in this setting. Structural nested models are based on counterfactuals: the outcome a person would have had had treatment been withheld after a certain time. Previous work on continuous-time structural nested models assumes that counterfactuals depend deterministically on observed data, while conjecturing that this assumption can be relaxed. This article proves that one can mimic counterfactuals by constructing random variables, solutions to a differential equation, that have the same distribution as the counterfactuals, even given past observed data. These "mimicking" variables can be used to estimate the parameters of structural nested models without assuming the treatment effect to be deterministic.

math.ST