SearcharxivSearch

arXiv subjects

Juraj Bodik

Publications and source records attributed to Juraj Bodik.

9 recordsLinked to original sources

Retrospective Counterfactual Prediction by Conditioning on the Factual Outcome: A Cross-World Approach

Retrospective causal questions ask what would have happened to an observed individual had they received a different treatment. We study the problem of estimating $\mu(x,y)=\mathbb{E}[Y(1)\mid X=x,Y(0)=y]$, the expected counterfactual outcome for an individual with covariates $x$ and observed outcome $y$, and constructing valid prediction intervals under the Neyman-Rubin superpopulation model. This quantity is generally not identified without additional assumptions. To link the observed and unobserved potential outcomes, we work with a cross-world correlation $\rho(x)=cor(Y(1),Y(0)\mid X=x)$; plausible bounds on $\rho(x)$ enable a principled approach to this otherwise unidentified problem. We introduce retrospective counterfactual estimators $\hat{\mu}_{\rho}(x,y)$ and prediction intervals $C_{\rho}(x,y)$ that asymptotically satisfy $P[Y(1)\in C_{\rho}(x,y)\mid X=x, Y(0)=y]\ge1-\alpha$ under standard causal assumptions. Many common baselines implicitly correspond to endpoint choices $\rho=0$ or $\rho=1$ (ignoring the factual outcome or treating the counterfactual as a shifted factual outcome). Interpolating between these cases through cross-world dependence yields substantial gains in both theory and practice.

stat.ME

Identifiability of causal graphs under nonadditive conditionally parametric causal models

Existing approaches to causal discovery often rely on restrictive modeling assumptions that limit their applicability in real-world settings, particularly when data are heavy-tailed or contain a mixture of discrete and continuous variables. Identifiability of causal graphs has been established under several structural models, including linear non-Gaussian models, post-nonlinear models, and location-scale models. However, these frameworks may not capture the diversity of distributions observed in practice. To address this, we introduce Conditionally Parametric Causal Models (CPCM), a flexible class of models where the conditional distribution of the effect, given its cause, belongs to a known parametric family such as Gaussian, Poisson, Gamma, or Pareto. These models are adaptable to a wide range of practical situations, where the cause influences not only the mean but also the variance or tail behavior of the effect. We demonstrate the identifiability of CPCM by leveraging the concept of sufficient statistics. Furthermore, we propose an algorithm for estimating the causal structure from random samples drawn from CPCM. We evaluate the empirical properties of our methodology on various datasets, demonstrating state-of-the-art performance across multiple benchmarks.

stat.ME

Cross-World Assumption and Refining Prediction Intervals for Individual Treatment Effects

While average treatment effects (ATE) and conditional average treatment effects (CATE) provide valuable population- and subgroup-level summaries, they fail to capture uncertainty at the individual level. For high-stakes decision-making, individual treatment effect (ITE) estimates must be accompanied by valid prediction intervals that reflect heterogeneity and unit-specific uncertainty. However, the fundamental unidentifiability of ITEs limits the ability to derive precise and reliable individual-level uncertainty estimates. To address this challenge, we investigate the role of a cross-world correlation parameter, $ \rho(x) = cor(Y(1), Y(0) | X = x) $, which describes the dependence between potential outcomes, given covariates, in the Neyman-Rubin super-population model with i.i.d. units. Although $ \rho $ is fundamentally unidentifiable, we argue that in most real-world applications, it is possible to impose reasonable and interpretable bounds informed by domain-expert knowledge. Given $\rho$, we design prediction intervals for ITE, achieving more stable and accurate coverage with substantially shorter widths; often less than 1/3 of those from competing methods. The resulting intervals satisfy coverage guarantees $P\big(Y(1) - Y(0) \in C_{ITE}(X)\big) \geq 1 - \alpha$ and are asymptotically optimal under Gaussian assumptions. We provide strong theoretical and empirical arguments that cross-world assumptions can make individual uncertainty quantification both practically informative and statistically valid.

stat.ME

CLEAR: Calibrated Learning for Epistemic and Aleatoric Risk

Accurate uncertainty quantification is critical for reliable predictive modeling. Existing methods typically address either aleatoric uncertainty due to measurement noise or epistemic uncertainty resulting from limited data, but not both in a balanced manner. We propose CLEAR, a calibration method with two distinct parameters, $\gamma_1$ and $\gamma_2$, to combine the two uncertainty components and improve the conditional coverage of predictive intervals for regression tasks. CLEAR is compatible with any pair of aleatoric and epistemic estimators; we show how it can be used with (i) quantile regression for aleatoric uncertainty and (ii) ensembles drawn from the Predictability-Computability-Stability (PCS) framework for epistemic uncertainty. Across 17 diverse real-world datasets, CLEAR achieves an average improvement of 28.3\% and 17.5\% in the interval width compared to the two individually calibrated baselines while maintaining nominal coverage. Similar improvements are observed when applying CLEAR to Deep Ensembles (epistemic) and Simultaneous Quantile Regression (aleatoric). The benefits are especially evident in scenarios dominated by high aleatoric or epistemic uncertainty. Project page: https://unco3892.github.io/clear/

stat.ML

Structural restrictions in local causal discovery: identifying direct causes of a target variable

We consider the problem of learning a set of direct causes of a target variable from an observational joint distribution. Learning directed acyclic graphs (DAGs) that represent the causal structure is a fundamental problem in science. Several results are known when the full DAG is identifiable from the distribution, such as assuming a nonlinear Gaussian data-generating process. Here, we are only interested in identifying the direct causes of one target variable (local causal structure), not the full DAG. This allows us to relax the identifiability assumptions and develop possibly faster and more robust algorithms. In contrast to the Invariance Causal Prediction framework, we only assume that we observe one environment without any interventions. We discuss different assumptions for the data-generating process of the target variable under which the set of direct causes is identifiable from the distribution. While doing so, we put essentially no assumptions on the variables other than the target variable. In addition to the novel identifiability results, we provide two practical algorithms for estimating the direct causes from a finite random sample and demonstrate their effectiveness on several benchmark and real datasets.

stat.ME

Granger Causality in Extremes

We introduce a rigorous mathematical framework for Granger causality in extremes, designed to identify causal links from extreme events in time series. Granger causality plays a pivotal role in uncovering directional relationships among time-varying variables. While this notion gains heightened importance during extreme and highly volatile periods, state-of-the-art methods primarily focus on causality within the body of the distribution, often overlooking causal mechanisms that manifest only during extreme events. Our framework is designed to infer causality mainly from extreme events by leveraging the causal tail coefficient. We establish equivalences between causality in extremes and other causal concepts, including (classical) Granger causality, Sims causality, and structural causality. We prove other key properties of Granger causality in extremes and show that the framework is especially helpful under the presence of hidden confounders. We also propose a novel inference method for detecting the presence of Granger causality in extremes from data. Our method is model-free, can handle non-linear and high-dimensional time series, outperforms current state-of-the-art methods in all considered setups, both in performance and speed, and was found to uncover coherent effects when applied to financial and extreme weather observations.

stat.ML

Extreme Treatment Effect: Extrapolating Dose-Response Function Into Extreme Treatment Domain

The potential outcomes framework serves as a fundamental tool for quantifying causal effects. The average dose-response function (also called the effect curve), denoted as (μ(t)), is typically of interest when dealing with a continuous treatment variable (exposure). The focus of this work is to determine the impact of an extreme level of treatment, potentially beyond the range of observed values--that is, estimating (μ(t)) for very large (t). Our approach is grounded in the field of statistics known as extreme value theory. We outline key assumptions for the identifiability of the extreme treatment effect. Additionally, we present a novel and consistent estimation procedure that can potentially reduce the dimension of the confounders to at most 3. This is a significant result since typically, the estimation of (μ(t)) is very challenging due to high-dimensional confounders. In practical applications, our framework proves valuable when assessing the effects of scenarios such as drug overdoses, extreme river discharges, or extremely high temperatures on a variable of interest.

stat.ME

Causality in extremes of time series

Consider two stationary time series with heavy-tailed marginal distributions. We aim to detect whether they have a causal relation, that is, if a change in one causes a change in the other. Usual methods for causal discovery are not well suited if the causal mechanisms only appear during extreme events. We propose a framework to detect a causal structure from the extremes of time series, providing a new tool to extract causal information from extreme events. We introduce the causal tail coefficient for time series, which can identify asymmetrical causal relations between extreme events under certain assumptions. This method can handle nonlinear relations and latent variables. Moreover, we mention how our method can help estimate a typical time difference between extreme events. Our methodology is especially well suited for large sample sizes, and we show the performance on the simulations. Finally, we apply our method to real-world space-weather and hydro-meteorological datasets.

math.ST

Detecting causal covariates for extreme dependence structures

Determining the causes of extreme events is a fundamental question in many scientific fields. An important aspect when modelling multivariate extremes is the tail dependence. In application, the extreme dependence structure may significantly depend on covariates. As for the general case of modelling including covariates, only some of the covariates are causal. In this paper, we propose a methodology to discover the causal covariates explaining the tail dependence structure between two variables. The proposed methodology for discovering causal variables is based on comparing observations from different environments or perturbations. It is a desired methodology for predicting extremal behaviour in a new, unobserved environment. The methodology is applied to a dataset of $\text{NO}_2$ concentration in the UK. Extreme $\text{NO}_2$ levels can cause severe health problems, and understanding the behaviour of concurrent severe levels is an important question. We focus on revealing causal predictors for the dependence between extreme $\text{NO}_2$ observations at different sites.

stat.ME