SearcharxivSearch

arXiv subjects

Dingke Tang

Publications and source records attributed to Dingke Tang.

5 recordsLinked to original sources

A Robust Framework for Two-Sample Mendelian Randomization under Population Heterogeneity

Mendelian randomization is a powerful tool for causal inference in observational studies. The two-sample summary-data design, which estimates genetic associations with exposures and outcomes in separate cohorts, is the most widely used Mendelian randomization approach in large-scale genomic studies. However, this approach relies on a strong assumption of population homogeneity across the two samples. In practice, available samples often differ in ancestry, demographics, socioeconomic factors, covariate adjustment, and measurement protocols. Violations of the homogeneity assumption can bias causal effect estimates and undermine the credibility of Mendelian randomization findings. We introduce a robust, model-free Mendelian randomization framework that directly addresses population heterogeneity in the two-sample summary-data setting. Our method avoids parametric assumptions about population differences and is designed to address real-world challenges, including measurement error, weak instruments, and pleiotropy. We show that the proposed estimator is consistent and asymptotically normal under heterogeneous designs, and may offer efficiency gains over the classic estimator even in homogeneous settings. Through numerical simulations and a real data analysis for estimating the causal effect of body mass index on high-density lipoprotein cholesterol across ancestrally diverse populations, we demonstrate the practical utility, stability, and robustness of our approach.

stat.ME

Average Treatment Effect Estimation with Non-binary Instrumental Variables

Non-binary instrumental variables, especially continuous ones, are common in practice. A binary recoding induces a Wald ratio but may discard useful variation and reduce efficiency. Although fully nonparametric approaches can in principle use the entire instrument, they often require high-dimensional nuisance estimation which can be unstable with rich covariates. We address this problem by developing a generalized Wald estimand for binary treatments that uses the full variation in a non-binary instrument. Under standard instrumental-variable assumptions and a homogeneity condition, the estimand yields a common identification formula for categorical and continuous instruments. We further develop its semiparametric efficiency theory and construct a locally efficient debiased estimator using risk-minimization reparameterizations and double cross-fitting to accommodate flexible machine learning while improving numerical stability. The central technical challenge is that the many Wald ratios generated by a non-binary instrument must agree, thereby imposing overidentifying restrictions on the observed-data law. In this setting, characterizing the tangent space is nonstandard: it requires a second-order parametric submodel, a construction that, to our knowledge, has not been standard in semiparametric efficiency theory. Simulations show stable performance across sample sizes and greater efficiency than estimators based on dichotomized instruments. In an application to the Princess Margaret Cancer Centre lung cancer cohort, associational analyses link excess body weight to lower two-year mortality, a seemingly protective pattern often called the obesity paradox. The proposed instrumental-variable analysis instead suggests increased mortality, pointing to residual confounding behind this paradox.

stat.ME

The synthetic instrument: From sparse association to sparse causation

In many observational studies, researchers are often interested in studying the effects of multiple exposures on a single outcome. Standard approaches for high-dimensional data such as the lasso assume the associations between the exposures and the outcome are sparse. These methods, however, do not estimate the causal effects in the presence of unmeasured confounding. In this paper, we consider an alternative approach that assumes the causal effects in view are sparse. We show that with sparse causation, the causal effects are identifiable even with unmeasured confounding. At the core of our proposal is a novel device, called the synthetic instrument, that in contrast to standard instrumental variables, can be constructed using the observed exposures directly. We show that under linear structural equation models, the problem of causal effect estimation can be formulated as an $\ell_0$-penalization problem, and hence can be solved efficiently using off-the-shelf software. Simulations show that our approach outperforms state-of-art methods in both low-dimensional and high-dimensional settings. We further illustrate our method using a mouse obesity dataset.

stat.ME

The Promises of Parallel Outcomes

A key challenge in causal inference from observational studies is the identification and estimation of causal effects in the presence of unmeasured confounding. In this paper, we introduce a novel approach for causal inference that leverages information in multiple outcomes to deal with unmeasured confounding. The key assumption in our approach is conditional independence among multiple outcomes. In contrast to existing proposals in the literature, the roles of multiple outcomes in our key identification assumption are symmetric, hence the name parallel outcomes. We show nonparametric identifiability with at least three parallel outcomes and provide parametric estimation tools under a set of linear structural equation models. Our proposal is evaluated through a set of synthetic and real data analyses.

stat.ME

Ultra-high Dimensional Variable Selection for Doubly Robust Causal Inference

Causal inference has been increasingly reliant on observational studies with rich covariate information. To build tractable causal procedures, such as the doubly robust estimators, it is imperative to first extract important features from high or even ultra-high dimensional data. In this paper, we propose causal ball screening for confounder selection from modern ultra-high dimensional data sets. Unlike the familiar task of variable selection for prediction modeling, our confounder selection procedure aims to control for confounding while improving efficiency in the resulting causal effect estimate. Previous empirical and theoretical studies suggest excluding causes of the treatment that are not confounders. Motivated by these results, our goal is to keep all the predictors of the outcome in both the propensity score and outcome regression models. A distinctive feature of our proposal is that we use an outcome model-free procedure for propensity score model selection, thereby maintaining double robustness in the resulting causal effect estimator. Our theoretical analyses show that the proposed procedure enjoys a number of properties, including model selection consistency and point-wise normality. Synthetic and real data analysis show that our proposal performs favorably with existing methods in a range of realistic settings. Data used in preparation of this article were obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database.

stat.ME