Searcharxiv⌕ Search

arXiv subjects

Iván Díaz

Publications and source records attributed to Iván Díaz.

At least 19 recordsLinked to original sources

Efficient transport and generalization of survival treatment effects

Randomized controlled trials provide internally valid estimates of treatment effects, but their results may not directly apply to broader target populations due to differences in baseline covariate distributions, adherence to treatment or variations in outcome mechanisms. Under standard transport and time-to-event identifiability assumptions, we develop nonparametric, debiased machine learning estimators for transporting and generalizing causal survival treatment effect differences from a source population to a target population in discrete time. We derive the efficient influence functions for the transport and generalization survival difference estimands and propose cross-fitted one-step estimators that are doubly robust and achieve semiparametric efficiency bounds under weak regularity conditions. We further introduce estimators that exploit known effect modifier subsets through an additive parameterization of the survival function, reducing the dimensionality of the reweighting and yielding smaller or equal asymptotic variance. We establish asymptotic normality, double robustness, and rates of convergence for all proposed estimators. Finite-sample properties are illustrated through Monte Carlo simulations under flexible and misspecified nuisance estimation scenarios. We apply the methods to data from the Women's Health Initiative to estimate the effect of hormone therapy on coronary heart disease across trial and observational populations.

stat.ME↗

Automated, efficient and model-free inference for randomized clinical trials via data-driven covariate adjustment

In 2023, the U.S. Food and Drug Administration issued guidance for adjustment of covariates in randomized clinical trials, emphasizing its role in enhancing precision and power through prognostic baseline variables. Despite its potential, many trials underutilize this method partly due to challenges in pre-specifying optimal baseline covariates and their functional forms. We explore the potential of automated, data-adaptive methods-including stepwise regression, Lasso and flexible machine learning algorithms-for covariate adjustment, addressing the challenge of pre-specification. Our approach ensures valid and interpretable treatment effect estimates and standard errors, even when outcome models are misspecified or biased outcome predictions are used. This differs from most competing methods, which assume correctly specified models for consistent standard errors. Our estimators require cross-fitting for reliable standard error estimation, though it can be omitted when variable selection is used, provided the outcome model satisfies an ultra-sparsity assumption. As such, we arrive at simple estimators and standard errors for marginal treatment effects in randomized clinical trials (or similar studies like A/B-testing), exploiting data-adaptive predictions from prognostic baseline covariates, with little (or no) bias in finite samples even when predictions are biased. Empirical and methodological results demonstrate promise of automated covariate adjustment for improving statistical power of trial analyses.

stat.ME↗

Modified treatment policies that depend on the natural history of treatment

Longitudinal modified treatment policies (LMTP) are a class of interventions that allow the definition, identification, and estimation of causal effects in general settings, such as with continuous or multivariate exposures, treatment regimens that require grace periods. Targeted machine learning estimators (i.e., double/debiased) have been formulated for LMTPs that assign the exposure at time $t$ as a function of the natural value of treatment at time $t$. However, important applications such as estimating the effect of a delay in the start of a treatment require formulating LMTPs that depend not only on the natural value of treatment at time $t$ but also on the \textit{history} of the natural value of treatment prior to time $t$. This paper develops targeted learning estimators for this general case. We discuss the definition of the effects, and propose estimators that use an augmented-data version of the sequential regression form of the longitudinal g-computation formula. Our estimators are based on the efficient influence function and provide $\sqrt{n}$ inference under standard doubly robust rate assumptions on the convergence of the outcome and treatment regressions. We apply the new estimators to assess the effect of delaying a risky pain treatment by one month on 12-month incidence of opioid use disorder.

stat.ME↗

Non-overlap Average Treatment Effect Bounds

The average treatment effect (ATE), the mean difference in potential outcomes under treatment and control, is a canonical causal effect. Overlap, which says that all subjects have non-zero probability of either treatment status, is necessary to identify and estimate the ATE. When overlap fails, the standard solution is to change the estimand, and target a trimmed effect in a subpopulation satisfying overlap. When the outcome is bounded, we demonstrate that this compromise is unnecessary. We derive non-overlap bounds: partial identification bounds on the ATE that do not require overlap. The bounds have width proportional to the size of the non-overlap subpopulation, making them informative in common scenarios when overlap violations are limited. Since the bounds are non-smooth functionals, we derive smooth approximations amenable to semiparametric efficiency theory and propose a Targeted Minimum Loss-Based estimator that is $\sqrt{n}$-consistent and asymptotically normal under nonparametric conditions. A multiplier bootstrap procedure yields uniformly valid confidence sets across all non-overlap subpopulation sizes and smoothing parameters, allowing researchers to report the tightest valid interval. Formally, we compare non-overlap confidence intervals to confidence intervals based on point estimation across multiple overlap regimes. We illustrate the method via simulation studies and real-world data applications.

stat.ME↗

Verification and Validation (V&V)-in-the-Loop for RISC-V Design: The Holistic Vision of BZL

The Barcelona Zetascale Lab (BZL) project aims to strengthening Europe's capacity in the design and manufacture of RISC-V based high-performance computing chips. In this context, we present a holistic pre-silicon verification and validation (V&V) methodology targeting highly robust RISC-V chip designs. This paper provides an overview of BZL's V&V approach, which integrates three complementary platforms: (1) a UVM-based verification environment to thoroughly validate RTL functionality; (2) an FPGA-based validation platform that enables system-level pre-silicon hardware-software RTL validation; and (3) a CI/CD flow that continuously automates build, deployment, and tests across these domains. By embedding these platforms into an industrial-grade V&V loop and exploiting large-scale CPU and FPGA hardware infrastructures, the BZL project enables continuous evolution of reliable hardware development and software integration. We believe that the BZL's V&V flow represents a robust and scalable foundation for ensuring the pre-silicon functional correctness and system level validation of RISC-V chip designs, and can serve as a key enabler for strategic initiatives in Europe, such as EPI and DARE, and beyond.

cs.AR↗

Orthogonal machine learning for conditional odds and risk ratios

Conditional effects are commonly used measures for understanding how treatment effects vary across different groups, and are often used to target treatments/interventions to groups who benefit most. In this work we review existing methods and propose novel ones, focusing on the odds ratio (OR) and the risk ratio (RR). While estimation of the conditional average treatment effect (ATE) has been widely studied, estimators for the OR and RR lag behind, and cutting edge estimators such as those based on doubly robust transformations or orthogonal risk functions have not been generalized to these parameters. We propose such a generalization here, focusing on the DR-learner and the R-learner. We derive orthogonal risk functions for the OR and RR and show that the associated pseudo-outcomes satisfy second-order conditional-mean remainder properties analogous to the ATE case. We also evaluate estimators for the conditional ATE, OR, and RR in a comprehensive nonparametric Monte Carlo simulation study to compare them with common alternatives under hundreds of different data-generating distributions. Our numerical studies provide empirical guidance for choosing an estimator. For instance, they show that while parametric models are useful in very simple settings, the proposed nonparametric estimators significantly reduce bias and mean squared error in the more complex settings expected in the real world. We illustrate the methods in the analysis of physical activity and sleep trouble in U.S. adults using data from the National Health and Nutrition Examination Survey (NHANES). The results demonstrate that our estimators uncover substantial treatment effect heterogeneity that is obscured by traditional regression approaches and lead to improved treatment decision rules, highlighting the importance of data-adaptive methods for advancing precision health research.

stat.ML↗

Comparing causal parameters with many treatments and positivity violations

Comparing outcomes across treatments is essential in medicine and public policy. To do so, researchers typically estimate a set of parameters, possibly counterfactual, with each targeting a different treatment. Treatment-specific means are commonly used, but their identification requires a positivity assumption, that every subject has a non-zero probability of receiving each treatment. This is often implausible, especially when treatment can take many values. Causal parameters based on dynamic stochastic interventions offer robustness to positivity violations. However, comparing these parameters may fail to reflect the effects of the underlying target treatments because the parameters can depend on outcomes under non-target treatments. To clarify when two parameters targeting different treatments yield a useful comparison of treatment efficacy, we propose a comparability criterion: if the conditional treatment-specific mean for one treatment is greater than that for another, then the corresponding causal parameter should also be greater. Many standard parameters fail to satisfy this criterion, but we show that only a mild positivity assumption is needed to identify parameters that yield useful comparisons. We then provide two simple examples that satisfy this criterion and are identifiable under the milder positivity assumption: trimmed and smooth trimmed treatment-specific means with multi-valued treatments. For smooth trimmed treatment-specific means, we develop doubly robust-style estimators that attain parametric convergence rates under nonparametric conditions. We illustrate our methods with an analysis of dialysis providers in New York State.

stat.ME↗

Computationally and statistically efficient estimation of time-smoothed counterfactual curves

Longitudinal causal inference is concerned with defining, identifying, and estimating the effect of a time-varying intervention on a time-varying outcome that is indexed by a follow-up time. In an observational study, Robins's generalized g-formula can identify causal effects induced by a broad class of time-varying interventions. Various methods for estimating the generalized g-formula have been posed for different outcome types, such as a failure event indicator by a specified time (e.g. mortality by 5 year follow-up), as well as continuous or dichotomous/multi-valued outcomes measures at a specified time (e.g. blood pressure in mm/hg or an indicator of high blood pressure at 5-year follow-up). Multiply-robust, data-adaptive estimators leverage flexible nonparametric estimation algorithms while allowing for statistical inference. However, extant methods do not accommodate time-smoothing when multiple outcomes are measured over time, which can lead to substantial loss of precision. We propose a novel multiply-robust estimator of the generalized g-formula that accommodates time-smoothing over numerous available outcome measures. Our approach accommodates any intervention that can be described as a Longitudinal Modified Treatment Policy, a flexible class suitable for binary, multi-valued, and continuous longitudinal treatments. Our method produces an estimate of the effect curve: the causal effect of the intervention on the outcome at each measurement time, taking into account censoring and non-monotonic outcome missingness patterns. In simulations we find that the proposed algorithm outperforms extant multiply-robust approaches for effect curve estimation in scenarios with high degrees of outcome missingness and when there is strong confounding. We apply the method to study longitudinal effects of union membership on wages.

stat.ME↗

Time-smoothed inverse probability weighted estimation of effects of generalized time-varying treatment strategies on repeated outcomes truncated by death

Researchers are often interested in estimating effects of generalized time-varying treatment strategies on the mean of an outcome at one or more selected follow-up times of interest. For example, the Medications and Weight Gain in PCORnet (MedWeight) study aimed to estimate effects of adhering to flexible medication regimes on future weight change using electronic health records (EHR) data. This problem presents several methodological challenges that have not been jointly addressed in the prior literature. First, this setting involves treatment strategies that vary over time and depend dynamically and non-deterministically on measured confounder history. Second, the outcome is repeatedly, non-monotonically, informatively, and sparsely measured in the data source. Third, some individuals die during follow-up, rendering the outcome of interest undefined at the follow-up time of interest. In this article, we pose a range of inverse probability weighted (IPW) estimators targeting effects of generalized time-varying treatment strategies in truncation by death settings that allow time-smoothing for precision gain. We conducted simulation studies that confirm precision gains of the time-smoothed IPW approaches over more conventional IPW approaches that do not leverage the repeated outcome measurements. We illustrate an application of the IPW approaches to estimate comparative effects of adhering to flexible antidepressant medication strategies on future weight change. The methods are implemented in the accompanying R package, smoothedIPW.

stat.ME↗

Identification and estimation of mediational effects of longitudinal modified treatment policies

We demonstrate a comprehensive semiparametric approach to causal mediation analysis, addressing the complexities inherent in settings with longitudinal and continuous treatments, confounders, and mediators. Our methodology utilizes a nonparametric structural equation model and a cross-fitted sequential regression technique based on doubly robust pseudo-outcomes, yielding an efficient, asymptotically normal estimator without relying on restrictive parametric modeling assumptions. We are motivated by a recent scientific controversy regarding the effects of invasive mechanical ventilation (IMV) on the survival of COVID-19 patients, considering acute kidney injury (AKI) as a mediating factor. We highlight the possibility of "inconsistent mediation," in which the direct and indirect effects of the exposure operate in opposite directions. We discuss the significance of mediation analysis for scientific understanding and its potential utility in treatment decisions.

stat.ME↗

Propensity score weighting across counterfactual worlds: longitudinal effects under positivity violations

When examining a contrast between two interventions, longitudinal causal inference studies frequently encounter positivity violations when one or both regimes are impossible to observe for some subjects. Existing weighting methods either assume positivity holds or produce effects that conflate interventions' impacts on ultimate outcomes with their effects on intermediate treatments and covariates. We propose a novel class of estimands -- cumulative cross-world weighted effects -- that weights potential outcome differences using propensity scores adapting to positivity violations cumulatively across timepoints and simultaneously across both counterfactual treatment histories. This new estimand isolates mechanistic differences between treatment regimes, is identifiable without positivity assumptions, and circumvents the limitations of existing longitudinal methods. Further, our analysis reveals two fundamental insights about longitudinal causal inference under positivity violations. First, while mechanistically meaningful, these effects correspond to non-implementable interventions, exposing a core interpretability-implementability tradeoff. Second, the identified effects faithfully capture mechanistic differences only under a partial common support assumption; violations cause the identified functional to collapse to zero, even when the causal effect is non-zero. We develop doubly robust-style estimators that achieve asymptotic normality and parametric convergence under nonparametric assumptions on the nuisance estimators. To this end, we reformulate challenging density ratio estimation as regression function estimation, which is achievable with standard machine learning methods. We illustrate our methods through analysis of union membership's effect on earnings.

stat.ME↗

General targeted machine learning for modern causal mediation analysis

Causal mediation analyses investigate the mechanisms through which causes exert their effects, and are therefore central to scientific progress. The literature on the non-parametric definition and identification of mediational effects in rigourous causal models has grown significantly in recent years, and there has been important progress to address challenges in the interpretation and identification of such effects. Despite great progress in the causal inference front, statistical methodology for non-parametric estimation has lagged behind, with few or no methods available for tackling non-parametric estimation in the presence of multiple, continuous, or high-dimensional mediators. In this paper we show that the identification formulas for six popular non-parametric approaches to mediation analysis proposed in recent years can be recovered from just two statistical estimands. We leverage this finding to propose an all-purpose one-step estimation algorithm that can be coupled with machine learning in any mediation study that uses any of these six definitions of mediation. The estimators have desirable properties, such as $\sqrt{n}$-convergence and asymptotic normality. Estimating the first-order correction for the one-step estimator requires estimation of complex density ratios on the potentially high-dimensional mediators, a challenge that is solved using recent advancements in so-called Riesz learning. We illustrate the properties of our methods in a simulation study and illustrate its use on real data to estimate the extent to which pain management practices mediate the total effect of having a chronic pain disorder on opioid use disorder.

stat.ML↗

Nonparametric estimation of an optimal treatment rule with fused randomized trials and missing effect modifiers

A fundamental principle of clinical medicine is that a treatment should only be administered to those patients who would benefit from it. Treatment strategies that assign treatment to patients as a function of their individual characteristics are known as dynamic treatment rules. The dynamic treatment rule that optimizes the outcome in the population is called the optimal dynamic treatment rule. Randomized clinical trials are considered the gold standard for estimating the marginal causal effect of a treatment on an outcome; they are often not powered to detect heterogeneous treatment effects, and thus, may rarely inform more personalized treatment decisions. The availability of multiple trials studying a common set of treatments presents an opportunity for combining data, often called data-fusion, to better estimate dynamic treatment rules. However, there may be a mismatch in the set of patient covariates measured across trials. We address this problem here; we propose a nonparametric estimator for the optimal dynamic treatment rule that leverages information across the set of randomized trials. We apply the estimator to fused randomized trials of medications for the treatment of opioid use disorder to estimate a treatment rule that would match patient subgroups with the medication that would minimize risk of return to regular opioid use.

stat.AP↗

Asymptotically Efficient Data-adaptive Penalized Shrinkage Estimation with Application to Causal Inference

A rich literature exists on constructing non-parametric estimators with optimal asymptotic properties. In addition to asymptotic guarantees, it is often of interest to design estimators with desirable finite-sample properties; such as reduced mean-squared error of a large set of parameters. We provide examples drawn from causal inference where this may be the case, such as estimating a large number of group-specific treatment effects. We show how finite-sample properties of non-parametric estimators, particularly their variance, can be improved by careful application of penalization. Given a target parameter of interest we derive a novel penalized parameter defined as the solution to an optimization problem that balances fidelity to the original parameter against a penalty term. By deriving the non-parametric efficiency bound for the penalized parameter, we are able to propose simple data-adaptive choices for the L1 and L2 tuning parameters designed to minimize finite-sample mean-squared error while preserving optimal asymptotic properties. The L1 and L2 penalization amounts to an adjustment that can be performed as a post-processing step applied to any asymptotically normal and efficient estimator. We show in extensive simulations that this adjustment yields estimators with lower MSE than the unpenalized estimators. Finally, we apply our approach to estimate provider quality measures of kidney dialysis providers within a causal inference framework.

stat.ME↗

Pulling back the curtain: the road from statistical estimand to machine-learning based estimator for epidemiologists (no wizard required)

Epidemiologists increasingly use causal inference methods that rely on machine learning, as these approaches can relax unnecessary model specification assumptions. While deriving and studying asymptotic properties of such estimators is a task usually associated with statisticians, it is useful for epidemiologists to understand the steps involved, as epidemiologists are often at the forefront of defining important new research questions and translating them into new parameters to be estimated. In this paper, our goal was to provide a relatively accessible guide through the process of (i) deriving an estimator based on the so-called efficient influence function (which we define and explain), and (ii) showing such an estimator's ability to validly incorporate machine learning, by demonstrating the so-called rate double robustness property. The derivations in this paper rely mainly on algebra and some foundational results from statistical inference, which are explained.

stat.ME↗

Causal survival analysis under competing risks using longitudinal modified treatment policies

Longitudinal modified treatment policies (LMTP) have been recently developed as a novel method to define and estimate causal parameters that depend on the natural value of treatment. LMTPs represent an important advancement in causal inference for longitudinal studies as they allow the non-parametric definition and estimation of the joint effect of multiple categorical, numerical, or continuous exposures measured at several time points. We extend the LMTP methodology to problems in which the outcome is a time-to-event variable subject to right-censoring and competing risks. We present identification results and non-parametric locally efficient estimators that use flexible data-adaptive regression techniques to alleviate model misspecification bias, while retaining important asymptotic properties such as $\sqrt{n}$-consistency. We present an application to the estimation of the effect of the time-to-intubation on acute kidney injury amongst COVID-19 hospitalized patients, where death by other causes is taken to be the competing event.

stat.ME↗

Longitudinal Generalizations of the Average Treatment Effect on the Treated for Multi-valued and Continuous Treatments

The Average Treatment Effect on the Treated (ATT) is a common causal parameter defined as the average effect of a binary treatment among the subset of the population receiving treatment. We propose a novel family of parameters, Generalized ATTs (GATTs), that generalize the concept of the ATT to longitudinal data structures, multi-valued or continuous treatments, and conditioning on arbitrary treatment subsets. We provide a formal causal identification result that expresses the GATT in terms of sequential regressions, and derive the efficient influence function of the parameter, which defines its semi-parametric efficiency bound. Efficient semi-parametric inference of the GATT requires estimating the ratios of functions of conditional probabilities (or densities); we propose directly estimating these ratios via empirical loss minimization, drawing on the theory of Riesz representers. Simulations suggest that estimation of the density ratios using Riesz representation have better stability in finite samples. Lastly, we illustrate the use of our methods to evaluate the effect of chronic pain management strategies on the development of opioid use disorder among Medicare patients with chronic pain.

stat.ME↗

Doubly Robust Nonparametric Efficient Estimation for Provider Evaluation

Provider profiling has the goal of identifying healthcare providers with exceptional patient outcomes. When evaluating providers, adjustment is necessary to control for differences in case-mix between different providers. Direct and indirect standardization are two popular risk adjustment methods. In causal terms, direct standardization examines a counterfactual in which the entire target population is treated by one provider. Indirect standardization, commonly expressed as a standardized outcome ratio, examines the counterfactual in which the population treated by a provider had instead been randomly assigned to another provider. Our first contribution is to present nonparametric efficiency bound for direct and indirectly standardized provider metrics by deriving their efficient influence functions. Our second contribution is to propose fully nonparametric estimators based on targeted minimum loss-based estimation that achieve the efficiency bounds. The finite-sample performance of the estimator is investigated through simulation studies. We apply our methods to evaluate dialysis facilities in New York State in terms of unplanned readmission rates using a large Medicare claims dataset. A software implementation of our methods is available in the R package TargetedRisk.

stat.ME↗