SearcharxivSearch

arXiv subjects

Davide Viviano

Publications and source records attributed to Davide Viviano.

16 recordsLinked to original sources

When is statistical evidence strong enough? Using hypothesis tests to value data collection

We recast statistical significance as a choice between making an immediate policy recommendation and deferring it until further evidence is collected. We show that the welfare-optimal decision corresponds, under minimax regret, to a statistical test whose level depends on the cost and precision of additional evidence. Inverting this rule, we introduce and recommend reporting the abstention-value (A-value) alongside traditional p-values to determine where additional data collection is most needed. The A-value defines the break-even welfare cost of abstaining and recommending further experimentation given the initial evidence. When experimentation capacity is limited, prioritizing additional data collection where A-values are the largest yields finite-sample welfare guarantees. We illustrate its implications for economic program evaluation.

econ.EM

Learning What to Learn: Experimental Design when Combining Experimental with Observational Evidence

Experiments deliver credible treatment-effect estimates but, because they are costly, are often restricted to specific sites, small populations, or particular mechanisms. A common practice across several fields is therefore to combine experimental estimates with reduced-form or structural external (observational) evidence to answer broader policy questions, such as those involving general equilibrium effects or external validity. We develop a unified framework for the design of experiments when combined with external evidence, i.e., choosing which experiment(s) to run and how to allocate sample size under arbitrary budget constraints. Because observational evidence may suffer bias unknown ex-ante, we evaluate designs using a robust regret criterion that compares any candidate design to an oracle with knowledge about the observational study bias bound that jointly chooses the design and estimator. This yields a transparent bias-variance trade-off that does not require the researcher to specify a bias bound and relies only on information already needed for conventional power calculations. We illustrate the framework for studying general equilibrium effects of cash transfer programs.

econ.EM

Policy design in experiments with unknown interference

This paper studies experimental designs for estimation and inference on policies with spillover effects. Units are organized into a finite number of large clusters and interact in unknown ways within each cluster. First, we introduce a single-wave experiment that, by varying the randomization across cluster pairs, estimates the marginal effect of a change in treatment probabilities, taking spillover effects into account. Using the marginal effect, we propose a test for policy optimality. Second, we design a multiple-wave experiment to estimate welfare-maximizing treatment rules. We provide strong theoretical guarantees and an implementation in a large-scale field experiment.

econ.EM

Evidence aggregation with ignorance in mind: learning what we do (not) know for archetypes discovery

When evaluating policy interventions, researchers often pursue two related goals: identifying which individuals or contexts benefit most, and determining whether patterns of treatment effect heterogeneity can be used to aggregate evidence across environments. We develop a framework that aggregates treatment effect heterogeneity, defined over individual and environmental characteristics, into interpretable summaries while setting aside contexts in which extrapolation is unreliable and further evidence is needed. The procedure therefore learns both how to summarize heterogeneous effects and when researchers should admit ignorance. We derive finite-sample regret guarantees, provide data-driven guarantees for selecting the complexity of the summary class, and inference procedures that quantify the value of follow-up data collection. We illustrate the approach by reanalyzing a multifaceted anti-poverty program implemented in six countries.

econ.EM

Experimental Design under Network Interference

This paper studies how to design two-wave experiments in the presence of spillovers for precise inference on treatment effects. We consider units connected through a single network, local dependence among individuals, and a general class of estimands encompassing average treatment and average spillover effects. We introduce a statistical framework for designing two-wave experiments with networks, where the researcher optimizes over participants and treatment assignments to minimize the variance of the estimators of interest, using a first-wave (pilot) experiment to estimate the variance. We derive guarantees for inference on treatment effects and regret guarantees on the variance obtained from the proposed design mechanism. Our results illustrate the existence of a trade-off in the choice of the pilot study and formally characterize the pilot's size relative to the main experiment. Simulations using simulated and real-world networks illustrate the advantages of the method.

econ.EM

Causal clustering: design of cluster experiments under network interference

This paper studies the design of cluster experiments to estimate the global treatment effect in the presence of network spillovers. We provide a framework to choose the clustering that minimizes the worst-case mean-squared error of the estimated global effect. We show that optimal clustering solves a novel penalized min-cut optimization problem computed via off-the-shelf semi-definite programming algorithms. Our analysis also characterizes simple conditions to choose between any two cluster designs, including choosing between a cluster or individual-level randomization. We illustrate the method's properties using unique network data from the universe of Facebook's users and existing data from a field experiment.

econ.EM

Program Evaluation with Remotely Sensed Outcomes

We study causal inference in experiments and quasi-experiments, where the economic outcome is imperfectly measured by a remotely sensed variable. The remotely sensed variable is low-cost, scalable, and predictive of the economic outcome in observational data; examples include satellite imagery and mobile phone activity. We model the remotely sensed variable as post-outcome: variation in the economic outcome causes variation in the remotely sensed variable. For example, changes in environmental quality cause changes in satellite imagery, not vice versa. Under this assumption, we propose a formula to nonparametrically identify the causal parameter by combining experimental and observational data. We develop a method for n^{-1/2} inference that is robust to misspecification and that does not restrict the algorithms used to process remotely sensed variables.

econ.EM

Estimating Social Norm Complementarities

We develop a model of choice over social norms that allows for complementarities along two dimensions: \textit{technological}, analogous to complementarities between consumption goods, and social, capturing returns from conformity. Together, these determine whether two norms are complements, substitutes, or independent, as defined by how the equilibrium prevalence of one norm responds to a marginal shift in the utility of another. We estimate the model using repeated cross-sections from Sierra Leone and Nigeria, focusing on female genital cutting, polygyny, and child marriage. Social returns are significant across all specifications. For female genital cutting and child marriage, we find evidence of complementarities, especially strong in Sierra Leone. For polygyny and child marriage, we find evidence of social substitutability, particularly in Nigeria. We interpret these differences using insights from anthropology. Finally, we iterate the model forward to study policy counterfactuals, assessing the potential effects of legal reforms and social interventions.

econ.GN

Dynamic covariate balancing: estimating treatment effects over time with potential local projections

This paper studies the estimation and inference of treatment effects in panel data settings when treatments change dynamically over time. We propose a balancing method that allows for (i) treatments to be assigned dynamically over time based on high-dimensional covariates, past outcomes, and treatments; (ii) outcomes and time-varying covariates to depend on the trajectory of all past treatments; (iii) heterogeneity of treatment effects. Our approach recursively projects potential outcomes' expectations on past histories. It then controls the bias arising from the non-experimental and sequential nature of this setting by balancing dynamically observable characteristics over time. We establish inferential guarantees of the proposed method even when the number of observable characteristics significantly exceeds the sample size. We study numerical properties of the estimator and illustrate the benefits of the procedure in an empirical application.

econ.EM

A model of multiple hypothesis testing

Multiple hypothesis testing practices vary widely, without consensus on which are appropriate when. This paper provides an economic foundation for these practices designed to capture leading examples, such as regulatory approval on the basis of clinical trials. MHT adjustments are appropriate in our framework to the extent that research costs are invariant to the number of hypotheses. Control of average size, as for example via a Bonferroni correction, emerges in the limit case where all costs are fixed; in the opposite limit, where costs vary in proportion to the hypothesis count, no correction is needed. We illustrate implications by calculating explicit critical values using data on actual costs in the drug approval process and in program evaluation research; these suggest that some MHT adjustment is warranted in these applications, but not as much as implied by standard practice.

econ.GN

Triply Robust Panel Estimators

This paper studies estimation of causal effects in a panel data setting. We introduce a new estimator, the Triply RObust Panel (TROP) estimator, that combines (i) a flexible model for the potential outcomes based on a low-rank factor structure on top of a two-way-fixed effect specification, with (ii) unit weights intended to upweight units similar to the treated units and (iii) time weights intended to upweight time periods close to the treated time periods. We study the performance of the estimator in a set of simulations designed to closely match several commonly studied real data sets. We find that there is substantial variation in the performance of the estimators across the settings considered. The proposed estimator outperforms two-way-fixed-effect/difference-in-differences, synthetic control, matrix completion and synthetic-difference-in-differences estimators. We investigate what features of the data generating process lead to this performance, and assess the relative importance of the three components of the proposed estimator. We have two recommendations. Our preferred strategy is that researchers use simulations closely matched to the data they are interested in, along the lines discussed in this paper, to investigate which estimators work well in their particular setting. A simpler approach is to use more robust estimators such as synthetic difference-in-differences or the new triply robust panel estimator which we find to substantially outperform two-way fixed effect estimators in many empirically relevant settings.

stat.ME

Publication Design with Incentives in Mind

The publication process both determines which research receives the most attention, and influences the supply of research through its impact on researchers' private incentives. We introduce a framework to study optimal publication decisions when researchers can choose (i) whether or how to conduct a study and (ii) whether or how to manipulate the research findings (e.g., via selective reporting or data manipulation). When manipulation is not possible, but research entails substantial private costs for the researchers, it may be optimal to incentivize cheaper research designs even if they are less accurate. When manipulation is possible, it is optimal to publish some manipulated results, as well as results that would have not received attention in the absence of manipulability. Even if it is possible to deter manipulation, such as by requiring pre-registered experiments instead of (potentially manipulable) observational studies, it is suboptimal to do so when experiments entail high research costs. We illustrate the implications of our model in an application to medical studies.

econ.EM

Policy Targeting under Network Interference

This paper studies the problem of optimally allocating treatments in the presence of spillover effects, using information from a (quasi-)experiment. I introduce a method that maximizes the sample analog of average social welfare when spillovers occur. I construct semi-parametric welfare estimators with known and unknown propensity scores and cast the optimization problem into a mixed-integer linear program, which can be solved using off-the-shelf algorithms. I derive a strong set of guarantees on regret, i.e., the difference between the maximum attainable welfare and the welfare evaluated at the estimated policy. The proposed method presents attractive features for applications: (i) it does not require network information of the target population; (ii) it exploits heterogeneity in treatment effects for targeting individuals; (iii) it does not rely on the correct specification of a particular structural model; and (iv) it accommodates constraints on the policy function. An application for targeting information on social networks illustrates the advantages of the method.

econ.EM

Identification and Inference for Synthetic Controls with Confounding

This paper studies inference on treatment effects in panel data settings with unobserved confounding. We model outcome variables through a factor model with random factors and loadings. Such factors and loadings may act as unobserved confounders: when the treatment is implemented depends on time-varying factors, and who receives the treatment depends on unit-level confounders. We study the identification of treatment effects and illustrate the presence of a trade-off between time and unit-level confounding. We provide asymptotic results for inference for several Synthetic Control estimators and show that different sources of randomness should be considered for inference, depending on the nature of confounding. We conclude with a comparison of Synthetic Control estimators with alternatives for factor models.

econ.EM

Synthetic learner: model-free inference on treatments over time

Understanding the effect of a particular treatment or a policy pertains to many areas of interest, ranging from political economics, marketing to healthcare. In this paper, we develop a non-parametric algorithm for detecting the effects of treatment over time in the context of Synthetic Controls. The method builds on counterfactual predictions from many algorithms without necessarily assuming that the algorithms correctly capture the model. We introduce an inferential procedure for detecting treatment effects and show that the testing procedure is asymptotically valid for stationary, beta mixing processes without imposing any restriction on the set of base algorithms under consideration. We discuss consistency guarantees for average treatment effect estimates and derive regret bounds for the proposed methodology. The class of algorithms may include Random Forest, Lasso, or any other machine-learning estimator. Numerical studies and an application illustrate the advantages of the method.

stat.ME

Fair Policy Targeting

One of the major concerns of targeting interventions on individuals in social welfare programs is discrimination: individualized treatments may induce disparities across sensitive attributes such as age, gender, or race. This paper addresses the question of the design of fair and efficient treatment allocation rules. We adopt the non-maleficence perspective of first do no harm: we select the fairest allocation within the Pareto frontier. We cast the optimization into a mixed-integer linear program formulation, which can be solved using off-the-shelf algorithms. We derive regret bounds on the unfairness of the estimated policy function and small sample guarantees on the Pareto frontier under general notions of fairness. Finally, we illustrate our method using an application from education economics.

econ.EM