SearcharxivSearch

arXiv subjects

Eric Auerbach

Publications and source records attributed to Eric Auerbach.

11 recordsLinked to original sources

Post-Selection Inference for Network Structure

Researchers often use the density of connections between groups of agents, such as communities, blocs, or markets, to characterize the structure of a social or economic network. In many cases, these groups are selected using the network data, making conventional fixed-group inference procedures potentially invalid. To address this issue, we develop two new confidence intervals that are universally valid post-selection in the sense that they guarantee simultaneous coverage asymptotically over all pairs of groups whose relative sizes do not vanish. Our first interval builds on a strategy of Berk et al. (2013). Our second interval is based on a Talagrand-type concentration inequality for empirical processes. Both intervals are simple to compute and scalable to large networks, but a key technical contribution of our paper is to show that the second interval is rate-optimal over a broader class of intervals. Three empirical illustrations show that accounting for selection can matter in practice. Some evidence for homophily in a social network and a hub-and-spoke structure in a trade network survives our correction, while evidence for a segmented market structure in a worker transition network does not.

econ.EM

Testing the Fairness-Accuracy Improvability of Algorithms

Many organizations use algorithms that have a disparate impact, i.e., the benefits or harms of the algorithm fall disproportionately on certain social groups. Addressing an algorithm's disparate impact can be challenging, however, because it is often unclear whether it is possible to reduce this impact without sacrificing other objectives of the organization, such as accuracy or profit. Establishing the improvability of algorithms with respect to multiple criteria is of both conceptual and practical interest: in many settings, disparate impact that would otherwise be prohibited under US federal law is permissible if it is necessary to achieve a legitimate business interest. The question is how a policy-maker can formally substantiate, or refute, this "necessity" defense. In this paper, we provide an econometric framework for testing the hypothesis that it is possible to improve on the fairness of an algorithm without compromising on other pre-specified objectives. Our proposed test is simple to implement and can be applied under any exogenous constraint on the algorithm space. We establish the large-sample validity and consistency of our test, and microfound the test's robustness to manipulation based on a game between a policymaker and the analyst. Finally, we apply our approach to evaluate a healthcare algorithm originally considered by Obermeyer et al. (2019), and quantify the extent to which the algorithm's disparate impact can be reduced without compromising the accuracy of its predictions.

econ.EM

Regression Discontinuity Design with Spillovers

This paper studies regression discontinuity designs (RDD) when linear-in-means spillovers occur between units that are close in their running variable. We show that the RDD estimand depends on the ratio of two terms: (1) the radius over which spillovers occur and (2) the choice of bandwidth used for the local linear regression. RDD estimates direct treatment effect when radius is of larger order than the bandwidth and total treatment effect when radius is of smaller order than the bandwidth. When the two are of similar order, the RDD estimand need not have a causal interpretation. To recover direct and spillover effects in the intermediate regime, we propose to incorporate estimated spillover terms into local linear regression. Our estimator is consistent and asymptotically normal and we provide bias-aware confidence intervals for direct treatment effects and spillovers. In the setting of Gonzalez (2021), we detect endogenous spillovers in voter fraud during the 2009 Afghan Presidential election. We also clarify when the donut-hole design addresses spillovers in RDD.

econ.EM

Exposure effects are not automatically useful for policymaking

We thank Savje (2023) for a thought-provoking article and appreciate the opportunity to share our perspective as social scientists. In his article, Savje recommends misspecified exposure effects as a way to avoid strong assumptions about interference when analyzing the results of an experiment. In this invited discussion, we highlight a limiation of Savje's recommendation: exposure effects are not generally useful for evaluating social policies without the strong assumptions that Savje seeks to avoid.

econ.EM

Identifying Socially Disruptive Policies

Social disruption occurs when a policy creates or destroys many network connections between agents. It is a costly side effect of many interventions and so a growing empirical literature recommends measuring and accounting for social disruption when evaluating the welfare impact of a policy. However, there is currently little work characterizing what can actually be learned about social disruption from data in practice. In this paper, we consider the problem of identifying social disruption in an experimental setting. We show that social disruption is not generally point identified, but informative bounds can be constructed by rearranging the eigenvalues of the marginal distribution of network connections between pairs of agents identified from the experiment. We apply our bounds to the setting of Banerjee et al. (2021) and find large disruptive effects that the authors miss by only considering regression estimates.

econ.EM

Heterogeneous Treatment Effects for Networks, Panels, and other Outcome Matrices

We are interested in the distribution of treatment effects for an experiment where units are randomized to a treatment but outcomes are measured for pairs of units. For example, we might measure risk sharing links between households enrolled in a microfinance program, employment relationships between workers and firms exposed to a trade shock, or bids from bidders to items assigned to an auction format. Such a double randomized experimental design may be appropriate when there are social interactions, market externalities, or other spillovers across units assigned to the same treatment. Or it may describe a natural or quasi experiment given to the researcher. In this paper, we propose a new empirical strategy that compares the eigenvalues of the outcome matrices associated with each treatment. Our proposal is based on a new matrix analog of the Fréchet-Hoeffding bounds that play a key role in the standard theory. We first use this result to bound the distribution of treatment effects. We then propose a new matrix analog of quantile treatment effects that is given by a difference in the eigenvalues. We call this analog spectral treatment effects.

econ.EM

Identification and Estimation of a Partially Linear Regression Model using Network Data

I study a regression model in which one covariate is an unknown function of a latent driver of link formation in a network. Rather than specify and fit a parametric network formation model, I introduce a new method based on matching pairs of agents with similar columns of the squared adjacency matrix, the ijth entry of which contains the number of other agents linked to both agents i and j. The intuition behind this approach is that for a large class of network formation models the columns of the squared adjacency matrix characterize all of the identifiable information about individual linking behavior. In this paper, I describe the model, formalize this intuition, and provide consistent estimators for the parameters of the regression model. Auerbach (2021) considers inference and an application to network peer effects.

econ.EM

Identification and Estimation of a Partially Linear Regression Model using Network Data: Inference and an Application to Network Peer Effects

This paper provides additional results relevant to the setting, model, and estimators of Auerbach (2019a). Section 1 contains results about the large sample properties of the estimators from Section 2 of Auerbach (2019a). Section 2 considers some extensions to the model. Section 3 provides an application to estimating network peer effects. Section 4 shows the results from some simulations.

econ.EM

The Local Approach to Causal Inference under Network Interference

We propose a new nonparametric modeling framework for causal inference when outcomes depend on how agents are linked in a social or economic network. Such network interference describes a large literature on treatment spillovers, social interactions, social learning, information diffusion, disease and financial contagion, social capital formation, and more. Our approach works by first characterizing how an agent is linked in the network using the configuration of other agents and connections nearby as measured by path distance. The impact of a policy or treatment assignment is then learned by pooling outcome data across similarly configured agents. We demonstrate the approach by deriving finite-sample bounds on the mean-squared error of a k-nearest-neighbor estimator for the average treatment response as well as proposing an asymptotically valid test for the hypothesis of policy irrelevance.

econ.EM

Testing for Differences in Stochastic Network Structure

How can one determine whether a community-level treatment, such as the introduction of a social program or trade shock, alters agents' incentives to form links in a network? This paper proposes analogues of a two-sample Kolmogorov-Smirnov test, widely used in the literature to test the null hypothesis of "no treatment effects", for network data. It first specifies a testing problem in which the null hypothesis is that two networks are drawn from the same random graph model. It then describes two randomization tests based on the magnitude of the difference between the networks' adjacency matrices as measured by the $2\to2$ and $\infty\to1$ operator norms. Power properties of the tests are examined analytically, in simulation, and through two real-world applications. A key finding is that the test based on the $\infty\to1$ norm can be substantially more powerful than that based on the $2\to2$ norm for the kinds of sparse and degree-heterogeneous networks common in economics.

econ.EM

Recovering Network Structure from Aggregated Relational Data using Penalized Regression

Social network data can be expensive to collect. Breza et al. (2017) propose aggregated relational data (ARD) as a low-cost substitute that can be used to recover the structure of a latent social network when it is generated by a specific parametric random effects model. Our main observation is that many economic network formation models produce networks that are effectively low-rank. As a consequence, network recovery from ARD is generally possible without parametric assumptions using a nuclear-norm penalized regression. We demonstrate how to implement this method and provide finite-sample bounds on the mean squared error of the resulting estimator for the distribution of network links. Computation takes seconds for samples with hundreds of observations. Easy-to-use code in R and Python can be found at https://github.com/mpleung/ARD.

econ.EM