SearcharxivSearch

arXiv subjects

Qingyuan Zhao

Publications and source records attributed to Qingyuan Zhao.

At least 19 recordsLinked to original sources

Counterfactual Optimization of Policy Interventions: Lexical Ordering and Leapfrogging

Most data-driven policy learning methods maximize average outcomes, overlooking the possibility that a policy beneficial on average may still harm a substantial fraction of individuals. Motivated by the ethical principle of "first do no harm", we study how to design a change from a baseline policy that improves overall welfare while keeping the worst-case probability or expectation of individual harm below a specified limit. We establish sufficient conditions under which an optimal policy transition has a lexical leapfrogging structure: groups defined by covariates and current treatment are ranked by a priority score, and any treatment change moves them directly to the conditionally optimal treatment. We derive this score under several models for the dependence among potential outcomes. We demonstrate this harm-aware policy optimization approach in a reanalysis of the I-SPY2 breast cancer platform trial and show how the consideration of counterfactual harm may lead to different conclusions about which treatment-subgroup pairs may warrant deprioritization in further clinical evaluation.

stat.ME

Apportioning Causal Responsibility of Two Risk Factors for an Adverse Outcome via Counterfactual Attribution

Unlike traditional causal inference, which prospectively evaluates the effects of causes, apportioning causal responsibility requires a retrospective assessment to deduce the causes of an outcome that has already occurred. This paper proposes a quantitative framework for apportioning causal responsibility between two binary risk factors that jointly contribute to a realized adverse outcome. Ideally, knowing the individual's latent causal type, defined by the potential outcomes under all possible exposure combinations, would allow precise apportionment; however, these potential outcomes cannot be simultaneously observed. We therefore define the average causal responsibility of each risk factor as its expected responsibility over the distribution of latent causal types. Under the assumptions of no confounding and monotonicity, we establish nonparametric identification of this metric when the type-specific responsibilities satisfy a structural balance condition, and derive sharp bounds otherwise. We illustrate the proposed framework using the classic example of lung cancer attributable to smoking and asbestos exposures.

stat.ME

An average-case sensitivity analysis for unmeasured confounding

Sensitivity analysis for the unconfoundedness assumption is crucial in observational studies. For this purpose, the marginal sensitivity model gained popularity recently due to good interpretability and mathematical properties. However, most existing models only consider a worst-case parameter that bounds the logit difference between the observed and full data propensity scores, which may not fully capture the extent of unmeasured confounding. We propose a new sensitivity model that is parameterized by the second moment of the propensity score ratio, requiring only the average strength of unmeasured confounding to be bounded. By characterizing the associated sensitivity analysis as an optimization problem, we derive sharp closed-form bounds of the average potential outcomes under our model. We propose efficient one-step estimators for these bounds based on the corresponding efficient influence functions. Additionally, we apply multiplier bootstrap to construct simultaneous confidence bands to cover the sensitivity curve that consists of bounds at different values of the sensitivity parameters. Through a real-data study, we illustrate how this average-case sensitivity analysis can provide tighter bounds and facilitate calibration of the results using observed covariates.

stat.ME

Selective Randomization Inference for Adaptive Experiments

Adaptive experiments use preliminary analyses of the data to inform further course of action and are commonly used in many disciplines including medical and social sciences. Because the null hypothesis and experimental design are data-dependent, it has long been recognized that statistical inference for adaptive experiments is not straightforward. Most existing methods only apply to specific adaptive designs and rely on strong assumptions. In this work, we propose selective randomization inference as a general framework for analysing adaptive experiments. In a nutshell, our approach applies conditional post-selection inference to randomization tests. By using directed acyclic graphs to describe the data generating process, we derive a selective randomization p-value that controls the selective type-I error. As inference only relies on the randomness in the treatment assignment, no modelling assumptions or independent and identically distributed data are needed. We elaborate on conditions that render the proposed p-value computable and provide rejection sampling and MCMC algorithms to find a Monte Carlo approximation. Moreover, this article shows how to estimate and construct confidence intervals for a homogeneous treatment effect. Lastly, we demonstrate our method and compare it with other randomization tests using synthetic and real-world data.

stat.ME

Heritability: A Counterfactual Perspective

Heritability is a central concept in the long-standing debate about nature versus nurture in biological and social sciences. However, existing notions of heritability are based on strong assumptions and do not use explicit causal models. We propose a new, counterfactual definition of heritability by adopting the potential outcomes model in causal inference. Our counterfactual heritability measures the importance of genetic inheritance by the average magnitude of difference between an individual with their hypothetical ``non-identical twin'' that is exposed to the exact same environment. We provide bounds on the counterfactual heritability that can, in principle, be computed from observational data. We then compare counterfactual heritability and its associated bounds with common notions of heritability in population-based studies, twin and sibling studies, and plant breeding experiments. Our results and comparisons highlight the importance of clarifying the causal structural assumptions and counterfactual comparisons in reasoning about heritability.

stat.AP

Simultaneous false discovery rate control in location families

When testing a number of statistical hypotheses using data from location families, it is often useful to control the false discovery rate (FDR) not just for hypotheses of the null values but also of other parameter values that are deemed practically insignificant. Here we consider FDR as a curve indexed by the location parameter and suggest a simple generalization of the Benjamini-Hochberg procedure that controls the FDR curve below any user-specified level. As a corollary of our main result, we show that the standard Benjamini-Hochberg procedure -- designed to control the FDR at the null -- also provides simultaneous control of the whole FDR curve for free. We further demonstrate the implications of our results and some practical considerations with a numerical example.

stat.ME

Counterfactual explainability and analysis of variance

Existing tools for explaining complex models and systems are associational rather than causal and do not provide mechanistic understanding. We propose a new notion called counterfactual explainability for causal attribution that is motivated by the concept of genetic heritability in twin studies. Counterfactual explainability extends methods for global sensitivity analysis (including the functional analysis of variance and Sobol's indices), which assumes independent explanatory variables, to dependent explanations by using a directed acyclic graphs to describe their causal relationship. Therefore, this explanability measure directly incorporates causal mechanisms by construction. Under a comonotonicity assumption, we discuss methods for estimating counterfactual explainability and apply them to a real dataset dataset to explain income inequality by gender, race, and educational attainment.

stat.ML

Confounder selection via iterative graph expansion

Confounder selection, namely choosing a set of covariates to control for confounding between a treatment and an outcome, is arguably the most important step in the design of an observational study. Previous methods, such as Pearl's back-door criterion, typically require pre-specifying a causal graph, which can often be difficult in practice. We propose an interactive procedure for confounder selection that does not require pre-specifying the graph or the set of observed variables. This procedure iteratively expands the causal graph by finding what we call "primary adjustment sets" for a pair of possibly confounded variables. This can be viewed as inverting a sequence of marginalizations of the underlying causal graph. Structural information in the form of primary adjustment sets is elicited from the user, bit by bit, until either a set of covariates is found to control for confounding or it can be determined that no such set exists. Other information, such as the causal relations between confounders, is not required by the procedure. We show that if the user correctly specifies the primary adjustment sets in every step, our procedure is both sound and complete.

stat.ME

Optimization-based Sensitivity Analysis for Unmeasured Confounding using Partial Correlations

Causal inference necessarily relies upon untestable assumptions; hence, it is crucial to assess the robustness of obtained results to violations of identification assumptions. However, such sensitivity analysis is only occasionally undertaken in practice, as many existing methods require analytically tractable solutions and their results are often difficult to interpret. We take a more flexible approach to sensitivity analysis and view it as a constrained stochastic optimization problem. This work focuses on sensitivity analysis for a linear causal effect when an unmeasured confounder and a potential instrument are present. We show how the bias of the OLS and TSLS estimands can be expressed in terms of partial correlations. Leveraging the algebraic rules that relate different partial correlations, practitioners can specify intuitive sensitivity models which bound the bias. We further show that the heuristic "plug-in" sensitivity interval may not have any confidence guarantees; instead, we propose a bootstrap approach to construct sensitivity intervals which performs well in numerical simulations. We illustrate the proposed methods with a real study on the causal effect of education on earnings and provide user-friendly visualization tools.

stat.ME

On statistical and causal models associated with acyclic directed mixed graphs

Causal models in statistics are often described using acyclic directed mixed graphs (ADMGs), which contain directed and bidirected edges and no directed cycles. This article surveys various interpretations of ADMGs, discusses their relations in different sub-classes of ADMGs, and argues that one of them -- the noise expansion (NE) model -- should be used as the default interpretation. Our endorsement of the NE model is based on two observations. First, in a subclass of ADMGs called unconfounded graphs (which retain most of the good properties of directed acyclic graphs and bidirected graphs), the NE model is equivalent to many other interpretations including the global Markov and nested Markov models. Second, the NE model for an arbitrary ADMG is exactly the union of that for all unconfounded expansions of that graph. This property is referred to as completeness, as it shows that the model does not commit to any specific latent variable explanation. In proving that the NE model is nested Markov, we also develop an ADMG-based theory for causality. Finally, we compare the NE model with the closely related but different interpretation of ADMGs as directed acyclic graphs (DAGs) with latent variables that is commonly used in the literature. We argue that the "latent DAG" interpretation is mathematically unnecessary, makes obscure ontological assumptions, and discourages practitioners from deliberating over important structural assumptions.

math.ST

Off-policy Evaluation with Deeply-abstracted States

Off-policy evaluation (OPE) is crucial for assessing a target policy's impact offline before its deployment. However, achieving accurate OPE in large state spaces remains challenging. This paper studies state abstractions -- originally designed for policy learning -- in the context of OPE. Our contributions are three-fold: (i) We define a set of irrelevance conditions central to learning state abstractions for OPE, and derive a backward-model-irrelevance condition for achieving irrelevance in %sequential and (marginalized) importance sampling ratios by constructing a time-reversed Markov decision process (MDP). (ii) We propose a novel iterative procedure that sequentially projects the original state space into a smaller space, resulting in a deeply-abstracted state, which substantially simplifies the sample complexity of OPE arising from high cardinality. (iii) We prove the Fisher consistencies of various OPE estimators when applied to our proposed abstract state spaces.

stat.ML

A Graphical Approach to State Variable Selection in Off-policy Learning

Sequential decision problems are widely studied across many areas of science. A key challenge when learning policies from historical data - a practice commonly referred to as off-policy learning - is how to ``identify'' the impact of a policy of interest when the observed data are not randomized. Off-policy learning has mainly been studied in two settings: dynamic treatment regimes (DTRs), where the focus is on controlling confounding in medical problems with short decision horizons, and offline reinforcement learning (RL), where the focus is on dimension reduction in closed systems such as games. The gap between these two well studied settings has limited the wider application of off-policy learning to many real-world problems. Using the theory for causal inference based on acyclic directed mixed graph (ADMGs), we provide a set of graphical identification criteria in general decision processes that encompass both DTRs and MDPs. We discuss how our results relate to the often implicit causal assumptions made in the DTR and RL literatures and further clarify several common misconceptions. Finally, we present a realistic simulation study for the dynamic pricing problem encountered in container logistics, and demonstrate how violations of our graphical criteria can lead to suboptimal policies.

stat.ME

A constructive approach to selective risk control

Many modern applications require using data to select the statistical tasks and make valid inference after selection. In this article, we provide a unifying approach to control for a class of selective risks. Our method is motivated by a reformulation of the celebrated Benjamini-Hochberg (BH) procedure for multiple hypothesis testing as the fixed point iteration of the Benjamini-Yekutieli (BY) procedure for constructing post-selection confidence intervals. Building on this observation, we propose a constructive approach to control extra-selection risk (where selection is made after decision) by iterating decision strategies that control the post-selection risk (where decision is made after selection). We show that many previous methods and results are special cases of this general framework, and we further extend this approach to problems with multiple selective risks. Our development leads to two surprising results about the BH procedure: (1) in the context of one-sided location testing, the BH procedure not only controls the false discovery rate at the null but also at other locations for free; (2) in the context of permutation tests, the BH procedure with exact permutation p-values can be well approximated by a procedure which only requires a total number of permutations that is almost linear in the total number of hypotheses.

stat.ME

Multiple conditional randomization tests for lagged and spillover treatment effects

We consider the problem of constructing multiple independent conditional randomization tests using a single dataset. Because the tests are independent, the randomization p-values can be interpreted individually and combined using standard methods for multiple testing. We give a simple, sequential construction of such tests, and then discuss its application to three problems: Rosenbaum's evidence factors for observational studies, lagged treatment effect in stepped-wedge trials, and spillover effect in randomized trials with interference. We compare the proposed approach with some existing methods using simulated and real datasets. Finally, we establish a more general sufficient condition for independent conditional randomization tests.

math.ST

Active Learning for Discovering Complex Phase Diagrams with Gaussian Processes

We introduce a Bayesian active learning algorithm that efficiently elucidates phase diagrams. Using a novel acquisition function that assesses both the impact and likelihood of the next observation, the algorithm iteratively determines the most informative next experiment to conduct and rapidly discerns the phase diagrams with multiple phases. Comparative studies against existing methods highlight the superior efficiency of our approach. We demonstrate the algorithm's practical application through the successful identification of the entire phase diagram of a spin Hamiltonian with antisymmetric interaction on Honeycomb lattice, using significantly fewer sample points than traditional grid search methods and a previous method based on support vector machines. Our algorithm identifies the phase diagram consisting of skyrmion, spiral and polarized phases with error less than 5% using only 8% of the total possible sample points, in both two-dimensional and three-dimensional phase spaces. Additionally, our method proves highly efficient in constructing three-dimensional phase diagrams, significantly reducing computational and experimental costs. Our methodological contributions extend to higher-dimensional phase diagrams with multiple phases, emphasizing the algorithm's effectiveness and versatility in handling complex, multi-phase systems in various dimensions.

physics.comp-ph

A matrix algebra for graphical statistical models

Directed mixed graphs permit directed and bidirected edges between any two vertices. They were first considered in the path analysis developed by Sewall Wright and play an essential role in statistical modeling. We introduce a matrix algebra for walks on such graphs. Each element of the algebra is a matrix whose entries are sets of walks on the graph from the corresponding row to the corresponding column. The matrix algebra is then generated by applying addition (set union), multiplication (concatenation), and transpose to the two basic matrices consisting of directed and bidirected edges. We use it to formalize, in the context of Gaussian linear systems, the correspondence between important graphical concepts such as latent projection and graph separation with important probabilistic concepts such as marginalization and (conditional) independence. In two further examples regarding confounder adjustment and the augmentation criterion, we illustrate how the algebra allows us to visualize complex graphical proofs. A "dictionary" and LATEX macros for the matrix algebra are provided in the Appendix.

math.ST

Confounder Selection: Objectives and Approaches

Confounder selection is perhaps the most important step in the design of observational studies. A number of criteria, often with different objectives and approaches, have been proposed, and their validity and practical value have been debated in the literature. Here, we provide a unified review of these criteria and the assumptions behind them. We list several objectives that confounder selection methods aim to achieve and discuss the amount of structural knowledge required by different approaches. Finally, we discuss limitations of the existing approaches and implications for practitioners.

stat.ME

Almost exact Mendelian randomization

Mendelian randomization (MR) is a natural experimental design based on the random transmission of genes from parents to offspring. However, this inferential basis is typically only implicit or used as an informal justification. As parent-offspring data becomes more widely available, we advocate a different approach to MR that is exactly based on this natural randomization, thereby formalizing the analogy between MR and randomized controlled trials. We begin by developing a causal graphical model for MR which represents several biological processes and phenomena, including population structure, gamete formation, fertilization, genetic linkage, and pleiotropy. This causal graph is then used to detect biases in population-based MR studies and identify sufficient confounder adjustment sets to correct these biases. We then propose a randomization test in the within-family MR design using the exogenous randomness in meiosis and fertilization, which is extensively studied in genetics. Besides its transparency and conceptual appeals, our approach also offers some practical advantages, including robustness to misspecified phenotype models, robustness to weak instruments, and elimination of bias arising from population structure, assortative mating, dynastic effects, and horizontal pleiotropy. We conclude with an analysis of a pair of negative and positive controls in the Avon Longitudinal Study of Parents and Children. The accompanying R package can be found at https://github.com/matt-tudball/almostexactmr.

stat.ME