SearcharxivSearch

arXiv subjects

Augusto Cerqua

Publications and source records attributed to Augusto Cerqua.

5 recordsLinked to original sources

Identifying Treatment and Spillover Effects with Control-Based and Forecast-Based Counterfactuals

Spillovers and interference pose fundamental challenges for causal inference, as treatment assigned to one unit may affect the outcome of others, violating the no-interference assumption underlying most empirical strategies. Existing approaches, based on partial interference, exposure mapping, spatial, network, or structural frameworks, typically rely on strong assumptions about interaction structures or require the existence of uncontaminated control units to estimate relevant causal parameters. We revisit this identification challenge within the potential outcomes framework and compare the conditions under which causal effects can be identified using two broad classes of counterfactual methods: control-based counterfactual methods (CBCMs), such as matching and difference-in-differences designs, and forecast-based counterfactual methods (FBCMs), including interrupted time-series and machine learning control methods. We show under which circumstances CBCMs and FBCMs identify average direct and spillover effects. Through simulations and an empirical application, we illustrate the main advantages and limitations of each approach. We show that, in the presence of pervasive or ill-defined spillover effects, CBCMs either cannot be used or entail severe identification concerns, whereas FBCMs can more credibly identify some of the causal parameters of interest, at least in the short term.

econ.EM

On the (Mis)Use of Machine Learning with Panel Data

We provide the first systematic assessment of data leakage issues in the use of machine learning on panel data. Our organizing framework clarifies why neglecting the cross-sectional and longitudinal structure of these data leads to hard-to-detect data leakage, inflated out-of-sample performance, and an inadvertent overestimation of the real-world usefulness and applicability of machine learning models. We then offer empirical guidelines for practitioners to ensure the correct implementation of supervised machine learning in panel data environments. An empirical application, using data from over 3,000 U.S. counties spanning 2000-2019 and focused on income prediction, illustrates the practical relevance of these points across nearly 500 models for both classification and regression tasks.

econ.EM

Heterogeneous Responses to Continuous Treatments: A Cluster-Based Causal Framework

When treatments are non-randomly assigned, continuous, and yield heterogeneous effects at the same intensity, causal identification becomes particularly challenging. In such contexts, existing approaches often fail to provide policy-relevant estimates of the relationship between treatment intensity and outcomes, especially in the presence of limited common support. To fill this gap, we introduce the Clustered Dose-Response Function (Cl-DRF), a novel estimator designed to uncover the continuous causal relationship between treatment intensity and the dependent variable across distinct subgroups. Our approach leverages both theoretical and data-driven sources of heterogeneity, relying on relaxed versions of the conditional independence and positivity assumptions that are plausible across various observational settings. We apply the Cl-DRF estimator to estimate subgroup-specific dose-response relationships between European Cohesion Funds and economic growth. In contrast to much of the literature, higher funding increases growth in more developed regions without diminishing returns, while limited absorptive capacity prevents other regions from fully benefiting.

econ.EM

Causal inference and policy evaluation without a control group

Without a control group, the most widespread methodologies for estimating causal effects cannot be applied. To fill this gap, we propose the Machine Learning Control Method, a new approach for causal panel analysis that estimates causal parameters without relying on untreated units. We formalize identification within the potential outcomes framework and then provide estimation based on machine learning algorithms. To illustrate the practical relevance of our method, we present simulation evidence, a replication study, and an empirical application on the impact of the COVID-19 crisis on educational inequality. We implement the proposed approach in the companion R package MachineControl

econ.EM

Was there a COVID-19 harvesting effect in Northern Italy?

We investigate the possibility of a harvesting effect, i.e. a temporary forward shift in mortality, associated with the COVID-19 pandemic by looking at the excess mortality trends of an area that registered one of the highest death tolls in the world during the first wave, Northern Italy. We do not find any evidence of a sizable COVID-19 harvesting effect, neither in the summer months after the slowdown of the first wave nor at the beginning of the second wave. According to our estimates, only a minor share of the total excess deaths detected in Northern Italian municipalities over the entire period under scrutiny (February - November 2020) can be attributed to an anticipatory role of COVID-19. A slightly higher share is detected for the most severely affected areas (the provinces of Bergamo and Brescia, in particular), but even in these territories, the harvesting effect can only account for less than 20% of excess deaths. Furthermore, the lower mortality rates observed in these areas at the beginning of the second wave may be due to several factors other than a harvesting effect, including behavioral change and some degree of temporary herd immunity. The very limited presence of short-run mortality displacement restates the case for containment policies aimed at minimizing the health impacts of the pandemic.

econ.GN