SearcharxivSearch

arXiv subjects

Bruno Ferman

Publications and source records attributed to Bruno Ferman.

16 recordsLinked to original sources

Treatment-effect heterogeneity and interactive fixed effects: Can we control for too much?

This paper studies the interactive fixed effects (IFE) estimator in a panel-data setting with heterogeneous treatment effects. We show that, if the treatment-effect heterogeneity admits a linear factor structure, the IFE estimator could fail to recover the average treatment effect on the treated units. The problem arises because the interactive fixed effects absorb the heterogeneity in the treatment effect, creating a \textit{bad-control} problem. With time-invariant factors or unit-invariant loadings in the treatment effect heterogeneity, identification may further break down due to multicollinearity. These problems are not present in alternative estimation methods that exclude treated units in post-treatment periods from the factor estimation.

econ.EM

On the Use of Design-Based Simulations

Design-based simulations - procedures that hold realized outcomes fixed and generate variation by resampling treatment assignment or shocks - are widely used in both methodological and applied work to assess inference procedures. This paper studies the extent to which such simulations are informative about inference validity. Focusing on shift-share designs, we show that standard simulations that fix outcomes and resample shocks may rely on a data-generating process that is not aligned with the true one. In particular, these simulations confound true treatment effects with error dependence, potentially overstating inference distortions due to spatial correlation. We propose alternative simulation designs that circumvent this problem and illustrate their use in prominent empirical applications. Our results highlight that the usefulness of design-based simulations depends critically on how closely the simulated data-generating process aligns with the true one.

econ.EM

Partial Identification under Stratified Randomization

This paper develops a unified framework for partial identification and inference in stratified experiments with attrition, accommodating both equal and heterogeneous treatment shares across strata. For equal-share designs, we apply recent theory for finely stratified experiments to Lee bounds, yielding closed-form, design-consistent variance estimators and properly sized confidence intervals. Simulations show that the conventional formula can overstate uncertainty, while our approach delivers tighter intervals. When treatment shares differ across strata, we propose a new strategy, which combines inverse probability weighting and global trimming to construct valid bounds even when strata are small or unbalanced. We establish identification, introduce a moment estimator, and extend existing inference results to stratified designs with heterogeneous shares, covering a broad class of moment-based estimators which includes the one we formulate. We also generalize our results to designs in which strata are defined solely by observed labels.

econ.EM

There must be an error here! Experimental evidence on coding errors' biases

Quantitative research relies heavily on coding, and coding errors are relatively common even in published research. In this paper, we examine whether individuals are more or less likely to check their code depending on the results they obtain. We test this hypothesis in a randomized experiment embedded in the recruitment process for research positions at a large international economic organization. In a coding task designed to assess candidates' programming abilities, we randomize whether participants obtain an expected or unexpected result if they commit a simple coding error. We find that individuals are almost 20% more likely to detect coding errors when they lead to unexpected results. This asymmetry in error detection depending on the results they generate suggests that coding errors may lead to biased findings in scientific research.

econ.GN

On the relationship between prediction intervals, tests of sharp nulls and inference on realized treatment effects in settings with few treated units

We study how inference methods for settings with few treated units that rely on treatment effect homogeneity extend to alternative inferential targets when treatment effects are heterogeneous -- namely, tests of sharp null hypotheses, inference on realized treatment effects, and prediction intervals. We show that inference methods for these alternative targets are deeply interconnected: they are either equivalent or become equivalent under additional assumptions. Our results show that methods designed under treatment effect homogeneity can remain valid for these alternative targets when treatment effects are stochastic, offering new theoretical justifications and insights on their applicability.

econ.EM

Inference with few treated units

In many causal inference applications, only one or a few units (or clusters of units) are treated. An important challenge in such settings is that standard inference methods relying on asymptotic theory may be unreliable, even with large total sample sizes. This survey reviews and categorizes inference methods designed to accommodate few treated units, considering cross-sectional and panel data methods. We discuss trade-offs and connections between different approaches. In doing so, we propose slight modifications to improve the finite-sample performance of some methods, and we also provide theoretical justifications for existing heuristic approaches that have been proposed in the literature.

econ.EM

Instrumental Variables with Time-Varying Exposure: Dynamic Effects of Revascularization on Quality of Life

This paper develops instrumental variables (IV) estimators for dynamic causal effects in randomized trials with imperfect compliance. These methods are applied to a randomized trial that assigned patients with ischemic heart disease to either an invasive treatment arm centered on revascularization or a control group meant to receive non-invasive medical therapy. As is common in such ``strategy trials,'' many participants assigned to treatment remained untreated while many assigned to control crossed over into treatment. Protocol non-compliance causes ITT estimates to diverge from the effect of treatment received, while conventional per-protocol analyses that condition on treatment received are compromised by selection bias. Extending the static potential-outcomes IV framework, the methods here identify average causal effects of treatment for dynamic compliers, the set of trial participants who comply with trial protocol at different follow-up horizons. IV estimates of revascularization effects on compliers' quality of life are markedly larger and more sustained than previously reported ITT and per-protocol estimates. We also show how to estimate average characteristics and marginal potential outcome means for dynamic compliers. These results are used to explain confounding in as-treated per-protocol estimates.

econ.EM

Dynamic LATEs with a Static Instrument

In many situations, researchers are interested in identifying dynamic effects of an irreversible treatment with a time-invariant binary instrumental variable (IV). For example, in evaluations of dynamic effects of training programs with a single lottery determining eligibility. A common approach in these situations is to report per-period IV estimates. Under a dynamic extension of standard IV assumptions, we show that such IV estimands identify a weighted sum of treatment effects for different latent groups and treatment exposures. However, there is possibility of negative weights. We discuss point and partial identification of dynamic treatment effects in this setting under different sets of assumptions.

econ.EM

Extensions for Inference in Difference-in-Differences with Few Treated Clusters

In settings with few treated units, Difference-in-Differences (DID) estimators are not consistent, and are not generally asymptotically normal. This poses relevant challenges for inference. While there are inference methods that are valid in these settings, some of these alternatives are not readily available when there is variation in treatment timing and heterogeneous treatment effects; or for deriving uniform confidence bands for event-study plots. We present alternatives in settings with few treated units that are valid with variation in treatment timing and/or that allow for uniform confidence bands.

econ.EM

Randomization Inference Tests for Shift-Share Designs

We consider the problem of inference in shift-share research designs. The choice between existing approaches that allow for unrestricted spatial correlation involves tradeoffs, varying in terms of their validity when there are relatively few or concentrated shocks, and in terms of the assumptions on the shock assignment process and treatment effects heterogeneity. We propose alternative randomization inference methods that combine the advantages of different approaches. These methods are valid in finite samples under relatively stronger assumptions, while asymptotically valid under weaker assumptions.

econ.EM

Inference in Difference-in-Differences with Few Treated Units and Spatial Correlation

We consider the problem of inference in Difference-in-Differences (DID) when there are few treated units and errors are spatially correlated. We first show that, when there is a single treated unit, some existing inference methods designed for settings with few treated and many control units remain asymptotically valid when errors are weakly dependent. However, these methods may be invalid with more than one treated unit. We propose a menu of alternatives that are asymptotically valid in this setting, even when the relevant distance metric across units is unavailable. These alternatives vary in terms of the length of the resulting confidence intervals and the strength of the required assumptions. Our methods are also valid for comparison-of-means estimators and for construction of prediction intervals for counterfactual imputation methods.

econ.EM

Assessing Inference Methods

We analyze different types of simulations that applied researchers can use to assess whether their inference methods reliably control false-positive rates. We show that different assessments involve trade-offs, varying in the types of problems they may detect, finite-sample performance, susceptibility to sequential-testing distortions, susceptibility to cherry-picking, and implementation complexity. We also show that a commonly used simulation to assess inference methods in shift-share designs can lead to misleading conclusions and propose alternatives. Overall, we provide novel insights and recommendations for applied researchers on how to choose, implement, and interpret inference assessments in their empirical applications.

econ.EM

Synthetic Controls with Imperfect Pre-Treatment Fit

We analyze the properties of the Synthetic Control (SC) and related estimators when the pre-treatment fit is imperfect. In this framework, we show that these estimators are generally biased if treatment assignment is correlated with unobserved confounders, even when the number of pre-treatment periods goes to infinity. Still, we show that a demeaned version of the SC method can substantially improve in terms of bias and variance relative to the difference-in-difference estimator. We also derive a specification test for the demeaned SC estimator in this setting with imperfect pre-treatment fit. Given our theoretical results, we provide practical guidance for applied researchers on how to justify the use of such estimators in empirical applications.

econ.EM

Matching Estimators with Few Treated and Many Control Observations

We analyze the properties of matching estimators when there are few treated, but many control observations. We show that, under standard assumptions, the nearest neighbor matching estimator for the average treatment effect on the treated is asymptotically unbiased in this framework. However, when the number of treated observations is fixed, the estimator is not consistent, and it is generally not asymptotically normal. Since standard inference methods are inadequate, we propose alternative inference methods, based on the theory of randomization tests under approximate symmetry, that are asymptotically valid in this framework. We show that these tests are valid under relatively strong assumptions when the number of treated observations is fixed, and under weaker assumptions when the number of treated observations increases, but at a lower rate relative to the number of control observations.

econ.EM

Inference in Difference-in-Differences: How Much Should We Trust in Independent Clusters?

We analyze the challenges for inference in difference-in-differences (DID) when there is spatial correlation. We present novel theoretical insights and empirical evidence on the settings in which ignoring spatial correlation should lead to more or less distortions in DID applications. We show that details such as the time frame used in the estimation, the choice of the treated and control groups, and the choice of the estimator, are key determinants of distortions due to spatial correlation. We also analyze the feasibility and trade-offs involved in a series of alternatives to take spatial correlation into account. Given that, we provide relevant recommendations for applied researchers on how to mitigate and assess the possibility of inference distortions due to spatial correlation.

econ.EM

On the Properties of the Synthetic Control Estimator with Many Periods and Many Controls

We consider the asymptotic properties of the Synthetic Control (SC) estimator when both the number of pre-treatment periods and control units are large. If potential outcomes follow a linear factor model, we provide conditions under which the factor loadings of the SC unit converge in probability to the factor loadings of the treated unit. This happens when there are weights diluted among an increasing number of control units such that a weighted average of the factor loadings of the control units asymptotically reconstructs the factor loadings of the treated unit. In this case, the SC estimator is asymptotically unbiased even when treatment assignment is correlated with time-varying unobservables. This result can be valid even when the number of control units is larger than the number of pre-treatment periods.

econ.EM