SearcharxivSearch

arXiv subjects

Bryan S. Graham

Publications and source records attributed to Bryan S. Graham.

12 recordsLinked to original sources

Moment Restrictions for Nonlinear Panel Data Models with Feedback

Many panel data methods, while allowing for general dependence between covariates and time-invariant agent-specific heterogeneity, place strong a priori restrictions on feedback: how past outcomes, covariates, and heterogeneity map into future covariate levels. Ruling out feedback entirely, as often occurs in practice, is unattractive in many dynamic economic settings. We provide a general characterization of all feedback and heterogeneity robust (FHR) moment conditions for nonlinear panel data models and present constructive methods to derive feasible moment-based estimators for specific models. We characterize semiparametric efficiency bounds in this case, quantifying the information loss associated with accommodating feedback as well as providing insight into how to construct estimators with good efficiency properties in practice. When FHR moments point-identify the target parameter - the case we emphasize - we show they attain the efficiency bound associated with the full model. These results apply both to the finite dimensional parameter indexing the parametric part of the model as well as to estimands that involve averages over the distribution of unobserved heterogeneity or features of the feedback process. Lastly, we illustrate our methods by providing a complete characterization of all FHR moment functions in the multi-spell mixed proportional hazards model, and compute efficient moment functions for both model parameters and average effects in this setting.

econ.EM

Identification in a Binary Choice Panel Data Model with a Predetermined Covariate

We study identification in a binary choice panel data model with a single \emph{predetermined} binary covariate (i.e., a covariate sequentially exogenous conditional on lagged outcomes and covariates). The choice model is indexed by a scalar parameter $θ$, whereas the distribution of unit-specific heterogeneity, as well as the feedback process that maps lagged outcomes into future covariate realizations, are left unrestricted. We provide a simple condition under which $θ$ is never point-identified, no matter the number of time periods available. This condition is satisfied in most models, including the logit one. We also characterize the identified set of $θ$ and show how to compute it using linear programming techniques. While $θ$ is not generally point-identified, its identified set is informative in the examples we analyze numerically, suggesting that meaningful learning about $θ$ may be possible even in short panels with feedback. As a complement, we report calculations of identified sets for an average partial effect, and find informative sets in this case as well.

econ.EM

Scenario Sampling for Large Supermodular Games

This paper introduces a simulation algorithm for evaluating the log-likelihood function of a large supermodular binary-action game. Covered examples include (certain types of) peer effect, technology adoption, strategic network formation, and multi-market entry games. More generally, the algorithm facilitates simulated maximum likelihood (SML) estimation of games with large numbers of players, $T$, and/or many binary actions per player, $M$ (e.g., games with tens of thousands of strategic actions, $TM=O(10^4)$). In such cases the likelihood of the observed pure strategy combination is typically (i) very small and (ii) a $TM$-fold integral who region of integration has a complicated geometry. Direct numerical integration, as well as accept-reject Monte Carlo integration, are computationally impractical in such settings. In contrast, we introduce a novel importance sampling algorithm which allows for accurate likelihood simulation with modest numbers of simulation draws.

econ.EM

An optimal test for strategic interaction in social and economic network formation between heterogeneous agents

Consider a setting where $N$ players, partitioned into $K$ observable types, form a directed network. Agents' preferences over the form of the network consist of an arbitrary network benefit function (e.g., agents may have preferences over their network centrality) and a private component which is additively separable in own links. This latter component allows for unobserved heterogeneity in the costs of sending and receiving links across agents (respectively out- and in- degree heterogeneity) as well as homophily/heterophily across the $K$ types of agents. In contrast, the network benefit function allows agents' preferences over links to vary with the presence or absence of links elsewhere in the network (and hence with the link formation behavior of their peers). In the null model which excludes the network benefit function, links form independently across dyads in the manner described by \cite{Charbonneau_EJ17}. Under the alternative there is interdependence across linking decisions (i.e., strategic interaction). We show how to test the null with power optimized in specific directions. These alternative directions include many common models of strategic network formation (e.g., "connections" models, "structural hole" models etc.). Our random utility specification induces an exponential family structure under the null which we exploit to construct a similar test which exactly controls size (despite the the null being a composite one with many nuisance parameters). We further show how to construct locally best tests for specific alternatives without making any assumptions about equilibrium selection. To make our tests feasible we introduce a new MCMC algorithm for simulating the null distributions of our test statistics.

econ.EM

Minimax Risk and Uniform Convergence Rates for Nonparametric Dyadic Regression

Let $i=1,\ldots,N$ index a simple random sample of units drawn from some large population. For each unit we observe the vector of regressors $X_{i}$ and, for each of the $N\left(N-1\right)$ ordered pairs of units, an outcome $Y_{ij}$. The outcomes $Y_{ij}$ and $Y_{kl}$ are independent if their indices are disjoint, but dependent otherwise (i.e., "dyadically dependent"). Let $W_{ij}=\left(X_{i}',X_{j}'\right)'$; using the sampled data we seek to construct a nonparametric estimate of the mean regression function $g\left(W_{ij}\right)\overset{def}{\equiv}\mathbb{E}\left[\left.Y_{ij}\right|X_{i},X_{j}\right].$ We present two sets of results. First, we calculate lower bounds on the minimax risk for estimating the regression function at (i) a point and (ii) under the infinity norm. Second, we calculate (i) pointwise and (ii) uniform convergence rates for the dyadic analog of the familiar Nadaraya-Watson (NW) kernel regression estimator. We show that the NW kernel regression estimator achieves the optimal rates suggested by our risk bounds when an appropriate bandwidth sequence is chosen. This optimal rate differs from the one available under iid data: the effective sample size is smaller and $d_W=\mathrm{dim}(W_{ij})$ influences the rate differently.

math.ST

Sparse network asymptotics for logistic regression

Consider a bipartite network where $N$ consumers choose to buy or not to buy $M$ different products. This paper considers the properties of the logistic regression of the $N\times M$ array of i-buys-j purchase decisions, $\left[Y_{ij}\right]_{1\leq i\leq N,1\leq j\leq M}$, onto known functions of consumer and product attributes under asymptotic sequences where (i) both $N$ and $M$ grow large and (ii) the average number of products purchased per consumer is finite in the limit. This latter assumption implies that the network of purchases is sparse: only a (very) small fraction of all possible purchases are actually made (concordant with many real-world settings). Under sparse network asymptotics, the first and last terms in an extended Hoeffding-type variance decomposition of the score of the logit composite log-likelihood are of equal order. In contrast, under dense network asymptotics, the last term is asymptotically negligible. Asymptotic normality of the logistic regression coefficients is shown using a martingale central limit theorem (CLT) for triangular arrays. Unlike in the dense case, the normality result derived here also holds under degeneracy of the network graphon. Relatedly, when there happens to be no dyadic dependence in the dataset in hand, it specializes to recently derived results on the behavior of logistic regression with rare events and iid data. Sparse network asymptotics may lead to better inference in practice since they suggest variance estimators which (i) incorporate additional sources of sampling variation and (ii) are valid under varying degrees of dyadic dependence.

econ.EM

Teacher-to-classroom assignment and student achievement

We study the effects of counterfactual teacher-to-classroom assignments on average student achievement in elementary and middle schools in the US. We use the Measures of Effective Teaching (MET) experiment to semiparametrically identify the average reallocation effects (AREs) of such assignments. Our findings suggest that changes in within-district teacher assignments could have appreciable effects on student achievement. Unlike policies which require hiring additional teachers (e.g., class-size reduction measures), or those aimed at changing the stock of teachers (e.g., VAM-guided teacher tenure policies), alternative teacher-to-classroom assignments are resource neutral; they raise student achievement through a more efficient deployment of existing teachers.

econ.EM

Network Data

Many economic activities are embedded in networks: sets of agents and the (often) rivalrous relationships connecting them to one another. Input sourcing by firms, interbank lending, scientific research, and job search are four examples, among many, of networked economic activities. Motivated by the premise that networks' structures are consequential, this chapter describes econometric methods for analyzing them. I emphasize (i) dyadic regression analysis incorporating unobserved agent-specific heterogeneity and supporting causal inference, (ii) techniques for estimating, and conducting inference on, summary network parameters (e.g., the degree distribution or transitivity index); and (iii) empirical models of strategic network formation admitting interdependencies in preferences. Current research challenges and open questions are also discussed.

econ.EM

Dyadic Regression

Dyadic data, where outcomes reflecting pairwise interaction among sampled units are of primary interest, arise frequently in social science research. Regression analyses with such data feature prominently in many research literatures (e.g., gravity models of trade). The dependence structure associated with dyadic data raises special estimation and, especially, inference issues. This chapter reviews currently available methods for (parametric) dyadic regression analysis and presents guidelines for empirical researchers.

econ.EM

Kernel Density Estimation for Undirected Dyadic Data

We study nonparametric estimation of density functions for undirected dyadic random variables (i.e., random variables defined for all n\overset{def}{\equiv}\tbinom{N}{2} unordered pairs of agents/nodes in a weighted network of order N). These random variables satisfy a local dependence property: any random variables in the network that share one or two indices may be dependent, while those sharing no indices in common are independent. In this setting, we show that density functions may be estimated by an application of the kernel estimation method of Rosenblatt (1956) and Parzen (1962). We suggest an estimate of their asymptotic variances inspired by a combination of (i) Newey's (1994) method of variance estimation for kernel estimators in the "monadic" setting and (ii) a variance estimator for the (estimated) density of a simple network first suggested by Holland and Leinhardt (1976). More unusual are the rates of convergence and asymptotic (normal) distributions of our dyadic density estimates. Specifically, we show that they converge at the same rate as the (unconditional) dyadic sample mean: the square root of the number, N, of nodes. This differs from the results for nonparametric estimation of densities and regression functions for monadic data, which generally have a slower rate of convergence than their corresponding sample mean.

math.ST

Testing for Externalities in Network Formation Using Simulation

We discuss a simplified version of the testing problem considered by Pelican and Graham (2019): testing for interdependencies in preferences over links among N (possibly heterogeneous) agents in a network. We describe an exact test which conditions on a sufficient statistic for the nuisance parameter characterizing any agent-level heterogeneity. Employing an algorithm due to Blitzstein and Diaconis (2011), we show how to simulate the null distribution of the test statistic in order to estimate critical values and/or p-values. We illustrate our methods using the Nyakatoke risk-sharing network. We find that the transitivity of the Nyakatoke network far exceeds what can be explained by degree heterogeneity across households alone.

stat.ME

Semiparametrically efficient estimation of the average linear regression function

Let Y be an outcome of interest, X a vector of treatment measures, and W a vector of pre-treatment control variables. Here X may include (combinations of) continuous, discrete, and/or non-mutually exclusive "treatments". Consider the linear regression of Y onto X in a subpopulation homogenous in W = w (formally a conditional linear predictor). Let b0(w) be the coefficient vector on X in this regression. We introduce a semiparametrically efficient estimate of the average beta0 = E[b0(W)]. When X is binary-valued (multi-valued) our procedure recovers the (a vector of) average treatment effect(s). When X is continuously-valued, or consists of multiple non-exclusive treatments, our estimand coincides with the average partial effect (APE) of X on Y when the underlying potential response function is linear in X, but otherwise heterogenous across agents. When the potential response function takes a general nonlinear/heterogenous form, and X is continuously-valued, our procedure recovers a weighted average of the gradient of this response across individuals and values of X. We provide a simple, and semiparametrically efficient, method of covariate adjustment for settings with complicated treatment regimes. Our method generalizes familiar methods of covariate adjustment used for program evaluation as well as methods of semiparametric regression (e.g., the partially linear regression model).

econ.EM