SearcharxivSearch

arXiv subjects

Nathan Canen

Publications and source records attributed to Nathan Canen.

11 recordsLinked to original sources

When Predictions Become Regressors: A Split-Sample Correction for Biases in Downstream Inference

Prediction-based methods, including Large Language Models (LLMs) and other machine learning techniques, are often used to construct measures of political phenomena that are difficult to quantify directly, such as policy positions in manifestos or emotions expressed on social media. In many applications, these prediction-generated measures are used as explanatory variables in regression models, even though they are measured with error. This leads to biased estimates. In this paper, we propose a simple solution to these biases: instrumental variables constructed from multiple measures created on independent splits of the original data. This approach is theoretically valid, easy to implement, and does not require new data. Through simulations, we show that this approach recovers estimates close to the true values, even in relatively small samples, while the standard approach can produce substantial bias in practice. We illustrate the method by revisiting two applications: whether gendered speech affects legislative outcomes in the German Parliament, and whether political risk influences poverty alleviation programs in China.

econ.EM

Empirical Challenges with Peers-of-Peers Instruments in the Linear-In-Means Model

In the linear-in-means model, endogeneity arises naturally due to the reflection problem. A common solution is to use Instrumental Variables (IVs) based on higher-order network links, such as using friends-of-friends' characteristics. In this paper, we show that such instruments are unlikely to work well in many applied settings due to a specific sparse/dense-network mechanism: in extremely sparse networks, friends-of-friends instruments may become degenerate, while in denser networks they may still provide too little first-stage information. This implies that the IVs may be weak or that the first-stage estimand is undefined. We use random graph theory to characterize the rates at which these issues arise for a benchmark class of random graphs. This allows us to link network topology to first-stage information accumulation and to identify when such instruments are likely to perform well. We show how existing weak-IV robust inference can be adapted to this environment, and how scaling the network provides an alternative specification that can mitigate some of these challenges. We provide extensive Monte Carlo simulations and revisit empirical applications, showing the prevalence of such issues in empirical practice, and how our results apply.

econ.EM

Revealed and Concealed Repression: Theory and Measurement

Regimes routinely conceal acts of repression. We show that observed repression may be negatively correlated with total repression, consisting of both revealed and concealed acts. This distortion can generate perverse effects for policy interventions designed to reduce repression and complicates inference about the causes and consequences of repression. We develop a model in which regimes choose whether to conceal repression and activists decide whether to challenge the regime. We identify two measurement problems - one due to concealment and one to deterrence. We construct indices of repression that account for these problems and show how these indices can be expressed in terms of observable variables by leveraging equilibrium relationships. We then propose an empirical strategy to estimate these indices. As a proof of concept, we apply this approach to Russia, estimating repression indices at a monthly frequency for 2020-2025.

econ.TH

Simple Inference on a Simplex-Valued Weight

In many applications, the parameter of interest involves a simplex-valued weight which is identified as a solution to an optimization problem. Examples include synthetic control methods with group-level weights and various methods of model averaging and forecast combinations. The simplex constraint on the weight poses a challenge in statistical inference due to the constraint potentially binding. In this paper, we propose a simple method of constructing a confidence set for the weight using an adaptive test based on the projection on a polyhedral cone and prove that the method is asymptotically uniformly valid. The procedure does not require tuning parameters or simulations to compute critical values. The confidence set accommodates both the cases of point-identification or set-identification of the weight. We illustrate the method with an empirical example.

econ.EM

Synthetic Decomposition for Counterfactual Predictions

Counterfactual predictions are challenging when the policy variable goes beyond its pre-policy support. However, in many cases, information about the policy of interest is available from different ("source") regions where a similar policy has already been implemented. In this paper, we propose a novel method of using such data from source regions to predict a new policy in a target region. Instead of relying on extrapolation of a structural relationship using a parametric specification, we formulate a transferability condition and construct a synthetic outcome-policy relationship such that it is as close as possible to meeting the condition. The synthetic relationship weighs both the similarity in distributions of observables and in structural relationships. We develop a general procedure to construct asymptotic confidence intervals for counterfactual predictions and prove its asymptotic validity. We then apply our proposal to predict average teenage employment in Texas following a counterfactual increase in the minimum wage.

econ.EM

Quantifying Theory in Politics: Identification, Interpretation and the Role of Structural Methods

The best empirical research in political science clearly defines substantive parameters of interest, presents a set of assumptions that guarantee its identification, and uses an appropriate estimator. We argue for the importance of explicitly integrating rigorous theory into this process and focus on the advantages of doing so. By integrating theoretical structure into one's empirical strategy, researchers can quantify the effects of competing mechanisms, consider the ex-ante effects of new policies, extrapolate findings to new environments, estimate model-specific theoretical parameters, evaluate the fit of a theoretical model, and test competing models that aim to explain the same phenomena. As a guide to such a methodology, we provide an overview of structural estimation, including formal definitions, implementation suggestions, examples, and comparisons to other methods.

stat.ME

Choosing The Best Incentives for Belief Elicitation with an Application to Political Protests

Many experiments elicit subjects' prior and posterior beliefs about a random variable to assess how information affects one's own actions. However, beliefs are multi-dimensional objects, and experimenters often only elicit a single response from subjects. In this paper, we discuss how the incentives offered by experimenters map subjects' true belief distributions to what profit-maximizing subjects respond in the elicitation task. In particular, we show how slightly different incentives may induce subjects to report the mean, mode, or median of their belief distribution. If beliefs are not symmetric and unimodal, then using an elicitation scheme that is mismatched with the research question may affect both the magnitude and the sign of identified effects, or may even make identification impossible. As an example, we revisit Cantoni et al.'s (2019) study of whether political protests are strategic complements or substitutes. We show that they elicit modal beliefs, while modal and mean beliefs may be updated in opposite directions following their experiment. Hence, the sign of their effects may change, allowing an alternative interpretation of their results.

econ.EM

Inference in Linear Dyadic Data Models with Network Spillovers

When using dyadic data (i.e., data indexed by pairs of units), researchers typically assume a linear model, estimate it using Ordinary Least Squares and conduct inference using ``dyadic-robust" variance estimators. The latter assumes that dyads are uncorrelated if they do not share a common unit (e.g., if the same individual is not present in both pairs of data). We show that this assumption does not hold in many empirical applications because indirect links may exist due to network connections, generating correlated outcomes. Hence, ``dyadic-robust'' estimators can be biased in such situations. We develop a consistent variance estimator for such contexts by leveraging results in network statistics. Our estimator has good finite sample properties in simulations, while allowing for decay in spillover effects. We illustrate our message with an application to politicians' voting behavior when they are seating neighbors in the European Parliament.

econ.EM

Counterfactual Analysis under Partial Identification Using Locally Robust Refinement

Structural models that admit multiple reduced forms, such as game-theoretic models with multiple equilibria, pose challenges in practice, especially when parameters are set-identified and the identified set is large. In such cases, researchers often choose to focus on a particular subset of equilibria for counterfactual analysis, but this choice can be hard to justify. This paper shows that some parameter values can be more "desirable" than others for counterfactual analysis, even if they are empirically equivalent given the data. In particular, within the identified set, some counterfactual predictions can exhibit more robustness than others, against local perturbations of the reduced forms (e.g. the equilibrium selection rule). We provide a representation of this subset which can be used to simplify the implementation. We illustrate our message using moment inequality models, and provide an empirical application based on a model with top-coded data.

econ.EM

A Decomposition Approach to Counterfactual Analysis in Game-Theoretic Models

Decomposition methods are often used for producing counterfactual predictions in non-strategic settings. When the outcome of interest arises from a game-theoretic setting where agents are better off by deviating from their strategies after a new policy, such predictions, despite their practical simplicity, are hard to justify. We present conditions in Bayesian games under which the decomposition-based predictions coincide with the equilibrium-based ones. In many games, such coincidence follows from an invariance condition for equilibrium selection rules. To illustrate our message, we revisit an empirical analysis in Ciliberto and Tamer (2009) on firms' entry decisions in the airline industry.

econ.EM

Estimating Local Interactions Among Many Agents Who Observe Their Neighbors

In various economic environments, people observe other people with whom they strategically interact. We can model such information-sharing relations as an information network, and the strategic interactions as a game on the network. When any two agents in the network are connected either directly or indirectly in a large network, empirical modeling using an equilibrium approach can be cumbersome, since the testable implications from an equilibrium generally involve all the players of the game, whereas a researcher's data set may contain only a fraction of these players in practice. This paper develops a tractable empirical model of linear interactions where each agent, after observing part of his neighbors' types, not knowing the full information network, uses best responses that are linear in his and other players' types that he observes, based on simple beliefs about the other players' strategies. We provide conditions on information networks and beliefs such that the best responses take an explicit form with multiple intuitive features. Furthermore, the best responses reveal how local payoff interdependence among agents is translated into local stochastic dependence of their actions, allowing the econometrician to perform asymptotic inference without having to observe all the players in the game or having to know the precise sampling process.

stat.ME