SearcharxivSearch

arXiv subjects

Hyunseung Kang

Publications and source records attributed to Hyunseung Kang.

At least 19 recordsLinked to original sources

Exact, Nonparametric Sensitivity Analysis for Observational Studies of Contingency Tables

In observational studies, contingency tables are commonly used to examine associations between categorical variables. However, any test of association in contingency tables may be biased by unmeasured confounding, and existing sensitivity analyses typically assume a binary treatment or impose strong parametric assumptions on non-binary treatments. We develop an exact (non-asymptotic) and nonparametric sensitivity analysis for unmeasured confounding in $I \times J$ and $I\times J\times K$ contingency tables, accommodating both non-binary treatments and outcomes. Extending Rosenbaum's generic bias sensitivity model, we derive a general method to compute the exact worst-case null distribution for any count-based test, including chi-squared and likelihood-based tests of association. We further provide specialized results for subfamilies of count-based tests that enable more efficient computation of the worst-case null distribution. Finally, we investigate test power in sensitivity analyses and show that tests exploiting all treatment and outcome levels achieve higher power than tests that dichotomize the categorical variables. We illustrate the proposed methods with a re-analysis of the effect of pre-kindergarten care on math achievement using data from the Early Childhood Longitudinal Study. The R package sensitivityIxJ that implements the method is on CRAN.

stat.ME

Towards Robust Matched Observational Studies with General Treatment Types: Consistency, Efficiency, and Adaptivity

To ensure reliable causal conclusions from observational studies, researchers routinely conduct sensitivity analysis to assess robustness to unmeasured confounding. In matched observational studies (one of the most popular observational study designs), two foundational concepts, design sensitivity and Bahadur-Rosenbaum efficiency, are used to quantify the robustness of test statistics and study designs in sensitivity analyses. Unfortunately, these measures of robustness are unavailable for non-binary treatments (e.g., continuous treatments) and consequently, prevailing recommendations about robust tests may be misleading. In this work, we provide a unified framework to quantify the robustness of test statistics and study designs for any treatment type. We first present a negative result about a popular, ad-hoc approach based on dichotomizing the treatment variable. Next, we consider two complementary parameterizations of the sensitivity parameter that apply to arbitrary treatment types and establish a one-to-one correspondence between them. Using these equivalent parameterizations, we generalize design sensitivity and Bahadur-Rosenbaum efficiency and derive unified formulas for both quantities under general treatment types. We also propose a general data-adaptive approach that combines candidate test statistics to enhance robustness against unmeasured confounding. Our results yield new insights into the robustness of tests and study designs in matched observational studies with non-binary treatments.

stat.ME

Transporting causal effects from a randomized trial without "transportability:" a case study of political advertising during U.S. elections

During the 2020 U.S. presidential election, Aggarwal et al. (2023) conducted a large-scale randomized experiment to evaluate a digital ad campaign against Trump in five battleground states. While the study found no effect on voter turnout, it's unclear whether this null result generalizes to other battleground states, notably Georgia, which played a unique role in the 2020 election and differs from the battleground states. Inspired by the study, we present a transfer learning framework to estimate treatment effects in a target population (e.g., Georgia) based on a randomized experiment from a source population (e.g., the five battleground states). Our framework is based on a sensitivity analysis that allows for violation of transportability, a popular yet impractical assumption which requires all differences between the source and target populations to be characterized by observed variables. Under our framework, we propose two estimators of the target treatment effect: a simple regression estimator with bootstrap, which we recommend for practitioners in this field, and an estimator based on the efficient influence function. Importantly, both estimators allow for covariates to differ between the target and the source populations, another common scenario in practice. We also propose a new, sample splitting approach to calibrate the sensitivity parameter. We apply our framework to estimate the effect of the ad campaign on voter turnout in Georgia during the 2020 election. Our findings indicate that small departures from transportability can lead to dramatically different ad effects across counties of Georgia. The direction of the effects is largely driven by racial composition: counties with higher White and lower Black percents tend to show positive effects, while counties with higher Latinx percents tend to show negative effects.

stat.AP

A Universal Framework for Factorial Matched Observational Studies with General Treatment Types: Design, Analysis, and Applications

Matching is one of the most widely used causal inference frameworks in observational studies. However, all the existing matching-based causal inference methods are designed for either a single treatment with general treatment types (e.g., binary, ordinal, or continuous) or factorial (multiple) treatments with binary treatments only. To our knowledge, no existing matching-based causal methods can handle factorial treatments with general treatment types. This critical gap substantially hinders the applicability of matching in many real-world problems, in which there are often multiple, potentially non-binary (e.g., continuous) treatment components. To address this critical gap, this work develops a universal framework for the design and analysis of factorial matched observational studies with general treatment types (e.g., binary, ordinal, or continuous). We first propose a two-stage non-bipartite matching algorithm that constructs matched sets of units with similar covariates but distinct combinations of treatment doses, thereby enabling valid estimation of both main and interaction effects. We then introduce a new class of generalized factorial Neyman-type estimands that provide model-free, finite-population-valid definitions of marginal and interaction causal effects under factorial treatments with general treatment types. Randomization-based Fisher-type and Neyman-type inference procedures are developed, including unbiased estimators, asymptotically valid variance estimators, and variance adjustments incorporating covariate information for improved efficiency. Finally, we illustrate the proposed framework through a county-level application that evaluates the causal impacts of work- and non-work-trip reductions (social distancing practices) on COVID-19-related and drug-related outcomes during the COVID-19 pandemic in the United States.

stat.ME

A Non-Bipartite Matching Framework for Difference-in-Differences with General Treatment Types

Difference-in-differences (DID) is one of the most widely used causal inference frameworks in observational studies. However, most existing DID methods are designed for binary treatments and cannot be readily applied to non-binary treatment settings. Although recent work has begun to extend DID to non-binary (e.g., continuous) treatments, these approaches typically require strong additional assumptions, including parametric outcome models or the presence of idealized comparison units with (nearly) static treatment levels over time (commonly called ``stayers'' or ``quasi-stayers''). In this technical note, we introduce a new non-bipartite matching framework for DID that naturally accommodates general treatment types (e.g., binary, ordinal, or continuous). Our framework makes three main contributions. First, we develop an optimal non-bipartite matching design for DID that jointly balances baseline covariates across comparable units (reducing bias) and maximizes contrasts in treatment trajectories over time (improving efficiency). Second, we establish a post-matching randomization condition, the design-based counterpart to the traditional parallel-trends assumption, which enables valid design-based inference. Third, we introduce the sample average DID ratio, a finite-population-valid and fully nonparametric causal estimand applicable to arbitrary treatment types. Our design-based approach that preserves the full treatment-dose information, avoids parametric assumptions, does not rely on the existence of stayers or quasi-stayers, and operates entirely within a finite-population framework, without appealing to hypothetical super-populations or outcome distributions.

stat.ME

SLOPE and Designing Robust Studies for Generalization

A popular task in generalization is to learn about a new, target population based on data from an existing, source population. This task relies on conditional exchangeability, which asserts that differences between the source and target populations are fully captured by observable characteristics of the two populations. Unfortunately, this assumption is often untenable in practice due to unobservable differences between the source and target populations. Worse, the assumption cannot be verified with data, warranting the need for robust data collection processes and study designs that are inherently less sensitive to violation of the assumption. In this paper, we propose SLOPE (Sensitivity of LOcal Perturbations from Exchangeability), a simple, intuitive, and novel measure that quantifies the sensitivity to local violation of conditional exchangeability. SLOPE combines ideas from sensitivity analysis in causal inference and derivative-based measure of robustness from Hampel (1974). Among other properties, SLOPE can help investigators to choose (a) a robust source or target population or (b) a robust estimand. Also, we show an analytic relationship between SLOPE and influence functions, which investigators can use to derive SLOPE given an influence function. We conclude with a re-analysis of a multi-national randomized experiment and illustrate the role of SLOPE in informing robust study designs for generalization.

stat.ME

Semiparametric Spatial Point Processes

We introduce a broad class of models called semiparametric spatial point process for making inference between spatial point patterns and spatial covariates. These models feature an intensity function with both parametric and nonparametric components. For the parametric component, we derive the semiparametric efficiency lower bound under Poisson point patterns and propose a point process double machine learning estimator that can achieve this lower bound. The proposed estimator for the parametric component is also shown to be consistent and asymptotically normal for non-Poisson point patterns. For the nonparametric component, we propose a kernel-based estimator and characterize its rates of convergence. Computationally, we introduce a fast, numerical approximation that transforms the proposed estimator into an estimator derived from weighted generalized partial linear models. We conclude with a simulation study and two real data analyses from ecology and hydrogeology.

stat.ME

A More Robust Approach to Multivariable Mendelian Randomization

Multivariable Mendelian randomization (MVMR) uses genetic variants as instrumental variables to infer the direct effects of multiple exposures on an outcome. However, unlike univariable Mendelian randomization, MVMR often faces greater challenges with many weak instruments, which can lead to bias not necessarily toward zero and inflation of type I errors. In this work, we introduce a new asymptotic regime that allows exposures to have varying degrees of instrument strength, providing a more accurate theoretical framework for studying MVMR estimators. Under this regime, our analysis of the widely used multivariable inverse-variance weighted method shows that it is often biased and tends to produce misleadingly narrow confidence intervals in the presence of many weak instruments. To address this, we propose a simple, closed-form modification to the multivariable inverse-variance weighted estimator to reduce bias from weak instruments, and additionally introduce a novel spectral regularization technique to improve finite-sample performance. We show that the resulting spectral-regularized estimator remains consistent and asymptotically normal under many weak instruments. Through simulations and real data applications, we demonstrate that our proposed estimator and asymptotic framework can enhance the robustness of MVMR analyses.

stat.ME

Estimating Treatment Effects using Multiple Surrogates: The Role of the Surrogate Score and the Surrogate Index

Estimating the long-term effects of treatments is of interest in many fields. A common challenge in estimating such treatment effects is that long-term outcomes are unobserved in the time frame needed to make policy decisions. One approach to overcome this missing data problem is to analyze treatments effects on an intermediate outcome, often called a statistical surrogate, if it satisfies the condition that treatment and outcome are independent conditional on the statistical surrogate. The validity of the surrogacy condition is often controversial. Here we exploit that fact that in modern datasets, researchers often observe a large number, possibly hundreds or thousands, of intermediate outcomes, thought to lie on or close to the causal chain between the treatment and the long-term outcome of interest. Even if none of the individual proxies satisfies the statistical surrogacy criterion by itself, using multiple proxies can be useful in causal inference. We focus primarily on a setting with two samples, an experimental sample containing data about the treatment indicator and the surrogates and an observational sample containing information about the surrogates and the primary outcome. We state assumptions under which the average treatment effect be identified and estimated with a high-dimensional vector of proxies that collectively satisfy the surrogacy assumption, and derive the bias from violations of the surrogacy assumption, and show that even if the primary outcome is also observed in the experimental sample, there is still information to be gained from using surrogates.

stat.ME

Identification and Inference with Invalid Instruments

Instrumental variables (IVs) are widely used to study the causal effect of an exposure on an outcome in the presence of unmeasured confounding. IVs require an instrument, a variable that is (A1) associated with the exposure, (A2) has no direct effect on the outcome except through the exposure, and (A3) is not related to unmeasured confounders. Unfortunately, finding variables that satisfy conditions (A2) or (A3) can be challenging in practice. This paper reviews works where instruments may not satisfy conditions (A2) or (A3), which we refer to as invalid instruments. We review identification and inference under different violations of (A2) or (A3), specifically under linear models, non-linear models, and heteroskedatic models. We conclude with an empirical comparison of various methods by re-analyzing the effect of body mass index on systolic blood pressure from the UK Biobank.

stat.ME

Minimum Resource Threshold Policy Under Partial Interference

When developing policies for prevention of infectious diseases, policymakers often set specific, outcome-oriented targets to achieve. For example, when developing a vaccine allocation policy, policymakers may want to distribute them so that at least a certain fraction of individuals in a census block are disease-free and spillover effects due to interference within blocks are accounted for. The paper proposes methods to estimate a block-level treatment policy that achieves a pre-defined, outcome-oriented target while accounting for spillover effects due to interference. Our policy, the minimum resource threshold policy (MRTP), suggests the minimum fraction of treated units required within a block to meet or exceed the target level of the outcome. We estimate the MRTP from empirical risk minimization using a novel, nonparametric, doubly robust loss function. We then characterize statistical properties of the estimated MRTP in terms of the excess risk bound. We apply our methodology to design a water, sanitation, and hygiene allocation policy for Senegal with the goal of increasing the proportion of households with no children experiencing diarrhea to a level exceeding a specified threshold. Our policy outperforms competing policies and offers new approaches to design allocation policies, especially in international development for communicable diseases.

stat.AP

A Groupwise Approach for Inferring Heterogeneous Treatment Effects in Causal Inference

Recently, there has been great interest in estimating the conditional average treatment effect using flexible machine learning methods. However, in practice, investigators often have working hypotheses about effect heterogeneity across pre-defined subgroups of study units, which we call the groupwise approach. The paper compares two modern ways to estimate groupwise treatment effects, a nonparametric approach and a semiparametric approach, with the goal of better informing practice. Specifically, we compare (a) the underlying assumptions, (b) efficiency and adaption to the underlying data generating models, and (c) a way to combine the two approaches. We also discuss how to test a key assumption concerning the semiparametric estimator and to obtain cluster-robust standard errors if study units in the same subgroups are correlated. We demonstrate our findings by conducting simulation studies and reanalyzing the Early Childhood Longitudinal Study.

stat.ME

Bayesian Causal Forests & the 2022 ACIC Data Challenge: Scalability and Sensitivity

We demonstrate how Hahn et al.'s Bayesian Causal Forests model (BCF) can be used to estimate conditional average treatment effects for the longitudinal dataset in the 2022 American Causal Inference Conference Data Challenge. Unfortunately, existing implementations of BCF do not scale to the size of the challenge data. Therefore, we developed flexBCF -- a more scalable and flexible implementation of BCF -- and used it in our challenge submission. We investigate the sensitivity of our results to the choice of propensity score estimation method and the use of sparsity-inducing regression tree priors. While we found that our overall point predictions were not especially sensitive to these modeling choices, we did observe that running BCF with flexibly estimated propensity scores often yielded better-calibrated uncertainty intervals.

stat.AP

Propensity Score Modeling: Key Challenges When Moving Beyond the No-Interference Assumption

The paper presents some models for the propensity score. Considerable attention is given to a recently popular, but relatively under-explored setting in causal inference where the no-interference assumption does not hold. We lay out some key challenges in propensity score modeling under interference and present a few promising models based on existing works on mixed effects models.

stat.ME

Semiparametric Efficient Dimension Reduction in multivariate regression with an Inner Envelope

Recently, Su and Cook proposed a dimension reduction technique called the inner envelope which can be substantially more efficient than the original envelope or existing dimension reduction techniques for multivariate regression. However, their technique relied on a linear model with normally distributed error, which may be violated in practice. In this work, we propose a semiparametric variant of the inner envelope that does not rely on the linear model nor the normality assumption. We show that our proposal leads to globally and locally efficient estimators of the inner envelope spaces. We also present a computationally tractable algorithm to estimate the inner envelope. Our simulations and real data analysis show that our method is both robust and efficient compared to existing dimension reduction methods in a diverse array of settings.

stat.ME

A More Efficient, Doubly Robust, Nonparametric Estimator of Treatment Effects in Multilevel Studies

When studying treatment effects in multilevel studies, investigators commonly use (semi-)parametric estimators, which make strong parametric assumptions about the outcome, the treatment, and/or the correlation structure between study units in a cluster. We propose a novel estimator of treatment effects that does not make such assumptions. Specifically, the new estimator is shown to be doubly robust, asymptotically Normal, and often more efficient than existing estimators, all without having to make any parametric modeling assumptions about the outcome, the treatment, and the correlation structure. We achieve this by estimating two non-standard nuisance functions in causal inference, the conditional propensity score and the outcome covariance model, using existing existing machine learning methods designed for independent and identically distributed (i.i.d) data. The new estimator is also demonstrated in simulated and real data where the new estimator is drastically more efficient than existing estimators, especially when studying cluster-specific treatment effects.

stat.ME

A Robust, Differentially Private Randomized Experiment for Evaluating Online Educational Programs With Sensitive Student Data

Randomized control trials (RCTs) have been the gold standard to evaluate the effectiveness of a program, policy, or treatment on an outcome of interest. However, many RCTs assume that study participants are willing to share their (potentially sensitive) data, specifically their response to treatment. This assumption, while trivial at first, is becoming difficult to satisfy in the modern era, especially in online settings where there are more regulations to protect individuals' data. The paper presents a new, simple experimental design that is differentially private, one of the strongest notions of data privacy. Also, using works on noncompliance in experimental psychology, we show that our design is robust against "adversarial" participants who may distrust investigators with their personal data and provide contaminated responses to intentionally bias the results of the experiment. Under our new design, we propose unbiased and asymptotically Normal estimators for the average treatment effect. We also present a doubly robust, covariate-adjusted estimator that uses pre-treatment covariates (if available) to improve efficiency. We conclude by using the proposed experimental design to evaluate the effectiveness of online statistics courses at the University of Wisconsin-Madison during the Spring 2021 semester, where many classes were online due to COVID-19.

stat.AP

Efficient Semiparametric Estimation of Network Treatment Effects Under Partial Interference

Recently, many estimators for network treatment effects have been proposed. But, their optimality properties in terms of semiparametric efficiency have yet to be resolved. We present a simple, yet flexible asymptotic framework to derive the efficient influence function and the semiparametric efficiency lower bound for a family of network causal effects under partial interference. An important corollary of our results is that one of the existing estimators by Liu et al. (2019) is locally efficient. We also present other estimators that are efficient and discuss results on adaptive estimation. We conclude by using the efficient estimators to study the direct and spillover effects of conditional cash transfer programs in Colombia.

stat.ME