SearcharxivSearch

arXiv subjects

Elie Tamer

Publications and source records attributed to Elie Tamer.

15 recordsLinked to original sources

Identification in Linear Quantile Panel Models

This paper studies identification in linear quantile panel models with unrestricted individual heterogeneity when the number of time periods is fixed and small. We impose strict exogeneity, whereby the conditional quantile restriction holds given the individual's complete regressor history and latent individual effect, but otherwise allow the disturbances to be arbitrarily dependent over time.

econ.EM

Nonparametric Bayesian Inference for Partially Identified Discrete Response Models

This paper proposes a nonparametric Bayesian inference framework for partially identified discrete response models. The key observation is that these models map a reduced-form conditional choice probability to an identified set. Consequently, nonparametric Bayesian inference for the conditional probability mass function leads to Bayesian inference for the identified set. The inference framework nests conditional moment inequalities and linear systems with unknown coefficients as special cases. Importantly, our proposal does not require converting conditional moments into unconditional moments or discretizing covariates. We show that the posterior is consistent for the true identified set when the model is correctly specified, show that the posterior can consistently detect model misspecification, and show posterior consistency for a pseudo-identified set that is valid under misspecification. We also verify the assumptions for a class of priors based on Gaussian processes that we use to implement our proposal. These priors offer similar flexibility to frequentist partial identification methods, and are computationally attractive because posterior sampling can be performed in closed-form. We also show that many of the ideas in this paper extend to continuous responses and aggregated discrete responses (e.g., market shares).

econ.EM

Stationary Errors and Quantile Regression in Short Panels

This paper studies a linear panel model with an unrestricted individual effect and a time- stationary idiosyncratic disturbance. We first show that stationarity is a strong restriction in a quantile model. In a linear conditional quantile specification with quantile-dependent slopes, equality of the conditional residual distributions across periods generically forces the slope coefficient to be constant over the quantile index. Thus, a stationary-error model identifies a common location coefficient rather than a collection of quantile-specific slope effects. We then develop a fixed-T estimator of this common coefficient. For each period, we run a cross- sectional quantile regression of the outcome on the full history of regressors. Stationarity makes the quantile projection of the composite individual effect and disturbance common across the period-specific regressions. Differences between diagonal and off-diagonal blocks of the resulting projection coefficients therefore identify the common slope whenever T>=2. We combine all such restrictions by a two-step minimum-distance estimator. The estimator is root-n-consistent and asymptotically normal with fixed T, permits unrestricted dependence across periods within an individual, and does not estimate the individual effects. We provide a consistent analytic covariance estimator, a cluster bootstrap, and an overidentification test of the projection restrictions implied by stationarity. Extensive Monte Carlo experiments show adequate performance under various designs.

econ.EM

Stochastic Potential Choices and Outcomes

Applied econometricians typically model each individual as having fixed outcomes under treatment and control and, in instrumental-variables (IV) settings, fixed treatment decisions under each value of the instrument. This paper asks what changes when outcomes and treatment allocations or choices are stochastic at the individual level. In the model, each individual has a stable (but possibly stochastic) response type consisting of two objects: a treatment choice probability under each state and a potential outcome distribution under each treatment-state pair. These stochastic potential outcomes change the interpretation of some familiar estimators. For instance, in the deterministic IV model, the estimand identifies treatment effect only for compliers-those whose treatment status switches with the instrument. Under stochastic treatment allocation or choice there is no such discrete subgroup: the estimand averages effects over the population, weighing each individual by how much the instrument, policy, or assignment rule moves their probability of treatment. The paper then gives an information-based foundation for stochastic choice, in which individuals act on expected gains given their information. Finally, repeated choices give the stable-kernel formulation empirical content: short panels identify moments of individual treatment probabilities, while long panels identify the joint dependence between treatment effects and the probability movements induced by an instrument or policy.

econ.EM

Partial Identification from LLM Prompts

Large language models are increasingly used as binary classifiers when the true label is latent. We study partial identification of the prevalence $\theta = P(X^* = 1)$ from panels of LLM reports whose errors may be arbitrarily dependent given the truth. The design of replication determines the observable, and hence the identifying content: repeated prompts to one model yield a count, several named models a response vector, and both a response matrix. Cast as a two-component finite mixture, the problem makes the identification failure transparent: absent restrictions that separate the latent components, the prevalence $\theta$ is completely unidentified, and weak stochastic-ordering restrictions (first-order dominance, monotone likelihood ratio, mean ordering) leave the identified set at $[0,1]$. Identifying power comes instead from externally calibrated scores and events, which discipline the mixture in the spirit of the misclassification and corrupted-data literature. We characterize the resulting bounds, establishing validity and sharpness, and give an exact account of the identifying information in the full score distribution beyond its mean. When named models are asked repeated versions of the same question, what identifies $\theta$ is not the number of positive answers but which models agree across prompts -- a feature a vote count discards. An extension derives implied bounds on regression coefficients when $X^*$ is a regressor of interest that is not directly observed.

econ.EM

Fast Online Inference on Semiparametric Models

This paper develops a framework for fast online inference on semiparametric models with large sample sizes and possibly many covariates. The computational algorithm itself is the object of statistical study: after a globally consistent warm start in the first phase, the path of averaged online iterates generated in the second phase automatically delivers estimators with optimal convergence rates and valid confidence sets. Both phases require only a single pass over the data stream and are well suited to streaming data or to settings with storage/privacy constraints. For semiparametric monotone index models, the averaged trajectory of the second phase lead to estimators that are automatically orthogonalized and satisfy the laws of the iterated logarithms, and policy functionals are updated along the same trajectory at negligible additional cost. The averaged trajectories satisfy functional central limit theorems, which yield fast online inference via random scaling and bypass the explicit variance estimation that complicates inference for semiparametric models. Applied to a fixed large sample, our online algorithm achieves substantial computational gains over corresponding offline procedures without sacrificing statistical performance. Monte Carlo experiments show adequate behavior. Our methods are applied to 19 million traffic-stop online records from the North Carolina State Patrol (Pierson et al. 2020) and to the international trade data of Helpman et al. (2008) with over 300 regressors. We also compare our estimator to its parametric benchmark in both empirical illustrations

econ.EM

Prediction Sets and Conformal Inference with Interval Outcomes

Given data on a random variable \(Y\), a prediction set with miscoverage level \(α\in (0,1)\) is a set that contains a new draw of \(Y\) with probability \(1-α\). Among all prediction sets satisfying this coverage property, the oracle prediction set is the one with minimal volume. The oracle prediction set offers a complementary view of the distribution of \(Y\), beyond point estimators such as the mean and quantiles, and has attracted considerable interest recently. This paper develops methods for estimating such prediction sets conditional on observed covariates when \(Y\) is \textit{censored} or \textit{interval-valued}. We characterise the oracle prediction set under partial identification induced by interval censoring and propose consistent estimators for both oracle prediction intervals and more general oracle prediction sets consisting of multiple disjoint intervals. In addition, we apply conformal inference to construct finite-sample valid prediction sets for interval outcomes that remain consistent as the sample size grows, using a conformity score tailored to interval data. The proposed procedure accounts for irreducible prediction uncertainty due to the stochastic nature of outcomes, modelling uncertainty arising from partial identification, and sampling uncertainty that vanishes as sample size increases. We conduct Monte Carlo simulations and two empirical applications using UK job postings data and the US Current Population Survey. The results demonstrate the robustness and efficiency of the proposed methods.

econ.EM

Inference on High Dimensional Selective Labeling Models

A class of simultaneous equation models arise in the many domains where observed binary outcomes are themselves a consequence of the existing choices of of one of the agents in the model. These models are gaining increasing interest in the computer science and machine learning literatures where they refer the potentially endogenous sample selection as the {\em selective labels} problem. Empirical settings for such models arise in fields as diverse as criminal justice, health care, and insurance. For important recent work in this area, see for example Lakkaruju et al. (2017), Kleinberg et al. (2018), and Coston et al.(2021) where the authors focus on judicial bail decisions, and where one observes the outcome of whether a defendant filed to return for their court appearance only if the judge in the case decides to release the defendant on bail. Identifying and estimating such models can be computationally challenging for two reasons. One is the nonconcavity of the bivariate likelihood function, and the other is the large number of covariates in each equation. Despite these challenges, in this paper we propose a novel distribution free estimation procedure that is computationally friendly in many covariates settings. The new method combines the semiparametric batched gradient descent algorithm introduced in Khan et al.(2023) with a novel sorting algorithms incorporated to control for selection bias. Asymptotic properties of the new procedure are established under increasing dimension conditions in both equations, and its finite sample properties are explored through a simulation study and an application using judicial bail data.

econ.EM

Heterogeneous Treatment Effects via Linear Dynamic Panel Data Models

We study the identification of heterogeneous, intertemporal treatment effects (TE) when potential outcomes depend on past treatments. First, applying a dynamic panel data model to observed outcomes, we show that an instrumental variable (IV) version of the estimand in Arellano and Bond (1991) recovers a non-convex (negatively weighted) aggregate of TE plus non-vanishing trends. We then provide conditions on sequential exchangeability (SE) of treatment and on TE heterogeneity that reduce such an IV estimand to a convex (positively weighted) aggregate of TE. Second, even when SE is generically violated, such estimands identify causal parameters when potential outcomes are generated by dynamic panel data models with some homogeneity or mild selection assumptions. Finally, we motivate SE and compare it with parallel trends (PT) in various settings with experimental data (when treatments are sequentially randomized) and observational data (when treatments are dynamic, rational choices under learning).

econ.EM

Counterfactual Analysis in Empirical Games

We address counterfactual analysis in empirical models of games with partially identified parameters, and multiple equilibria and/or randomized strategies, by constructing and analyzing the counterfactual predictive distribution set (CPDS). This framework accommodates various outcomes of interest, including behavioral and welfare outcomes. It allows a variety of changes to the environment to generate the counterfactual, including modifications of the utility functions, the distribution of utility determinants, the number of decision makers, and the solution concept. We use a Bayesian approach to summarize statistical uncertainty. We establish conditions under which the population CPDS is sharp from the point of view of identification. We also establish conditions under which the posterior CPDS is consistent if the posterior distribution for the underlying model parameter is consistent. Consequently, our results can be employed to conduct counterfactual analysis after a preliminary step of identifying and estimating the underlying model parameter based on the existing literature. Our consistency results involve the development of a new general theory for Bayesian consistency of posterior distributions for mappings of sets. Although we primarily focus on a model of a strategic game, our approach is applicable to other structural models with similar features.

econ.EM

Parallel Trends and Dynamic Choices

Difference-in-differences is a common method for estimating treatment effects, and the parallel trends condition is its main identifying assumption: the trend in mean untreated outcomes is independent of the observed treatment status. In observational settings, treatment is often a dynamic choice made or influenced by rational actors, such as policy-makers, firms, or individual agents. This paper relates parallel trends to economic models of dynamic choice. We clarify the implications of parallel trends on agent behavior and study when dynamic selection motives lead to violations of parallel trends. Finally, we consider identification under alternative assumptions that accommodate features of dynamic choice.

econ.EM

Estimating High Dimensional Monotone Index Models by Iterative Convex Optimization1

In this paper we propose new approaches to estimating large dimensional monotone index models. This class of models has been popular in the applied and theoretical econometrics literatures as it includes discrete choice, nonparametric transformation, and duration models. A main advantage of our approach is computational. For instance, rank estimation procedures such as those proposed in Han (1987) and Cavanagh and Sherman (1998) that optimize a nonsmooth, non convex objective function are difficult to use with more than a few regressors and so limits their use in with economic data sets. For such monotone index models with increasing dimension, we propose to use a new class of estimators based on batched gradient descent (BGD) involving nonparametric methods such as kernel estimation or sieve estimation, and study their asymptotic properties. The BGD algorithm uses an iterative procedure where the key step exploits a strictly convex objective function, resulting in computational advantages. A contribution of our approach is that our model is large dimensional and semiparametric and so does not require the use of parametric distributional assumptions.

econ.EM

Efficient Estimation in NPIV Models: A Comparison of Various Neural Networks-Based Estimators

Artificial Neural Networks (ANNs) can be viewed as nonlinear sieves that can approximate complex functions of high dimensional variables more effectively than linear sieves. We investigate the performance of various ANNs in nonparametric instrumental variables (NPIV) models of moderately high dimensional covariates that are relevant to empirical economics. We present two efficient procedures for estimation and inference on a weighted average derivative (WAD): an orthogonalized plug-in with optimally-weighted sieve minimum distance (OP-OSMD) procedure and a sieve efficient score (ES) procedure. Both estimators for WAD use ANN sieves to approximate the unknown NPIV function and are root-n asymptotically normal and first-order equivalent. We provide a detailed practitioner's recipe for implementing both efficient procedures. We compare their finite-sample performances in various simulation designs that involve smooth NPIV function of up to 13 continuous covariates, different nonlinearities and covariate correlations. Some Monte Carlo findings include: 1) tuning and optimization are more delicate in ANN estimation; 2) given proper tuning, both ANN estimators with various architectures can perform well; 3) easier to tune ANN OP-OSMD estimators than ANN ES estimators; 4) stable inferences are more difficult to achieve with ANN (than spline) estimators; 5) there are gaps between current implementations and approximation theories. Finally, we apply ANN NPIV to estimate average partial derivatives in two empirical demand examples with multivariate covariates.

econ.EM

Inference on Auctions with Weak Assumptions on Information

Given a sample of bids from independent auctions, this paper examines the question of inference on auction fundamentals (e.g. valuation distributions, welfare measures) under weak assumptions on information structure. The question is important as it allows us to learn about the valuation distribution in a robust way, i.e., without assuming that a particular information structure holds across observations. We leverage the recent contributions of \cite{Bergemann2013} in the robust mechanism design literature that exploit the link between Bayesian Correlated Equilibria and Bayesian Nash Equilibria in incomplete information games to construct an econometrics framework for learning about auction fundamentals using observed data on bids. We showcase our construction of identified sets in private value and common value auctions. Our approach for constructing these sets inherits the computational simplicity of solving for correlated equilibria: checking whether a particular valuation distribution belongs to the identified set is as simple as determining whether a {\it linear} program is feasible. A similar linear program can be used to construct the identified set on various welfare measures and counterfactual objects. For inference and to summarize statistical uncertainty, we propose novel finite sample methods using tail inequalities that are used to construct confidence regions on sets. We also highlight methods based on Bayesian bootstrap and subsampling. A set of Monte Carlo experiments show adequate finite sample properties of our inference procedures. We illustrate our methods using data from OCS auctions.

econ.EM

Monte Carlo Confidence Sets for Identified Sets

In complicated/nonlinear parametric models, it is generally hard to know whether the model parameters are point identified. We provide computationally attractive procedures to construct confidence sets (CSs) for identified sets of full parameters and of subvectors in models defined through a likelihood or a vector of moment equalities or inequalities. These CSs are based on level sets of optimal sample criterion functions (such as likelihood or optimally-weighted or continuously-updated GMM criterions). The level sets are constructed using cutoffs that are computed via Monte Carlo (MC) simulations directly from the quasi-posterior distributions of the criterions. We establish new Bernstein-von Mises (or Bayesian Wilks) type theorems for the quasi-posterior distributions of the quasi-likelihood ratio (QLR) and profile QLR in partially-identified regular models and some non-regular models. These results imply that our MC CSs have exact asymptotic frequentist coverage for identified sets of full parameters and of subvectors in partially-identified regular models, and have valid but potentially conservative coverage in models with reduced-form parameters on the boundary. Our MC CSs for identified sets of subvectors are shown to have exact asymptotic coverage in models with singularities. We also provide results on uniform validity of our CSs over classes of DGPs that include point and partially identified models. We demonstrate good finite-sample coverage properties of our procedures in two simulation experiments. Finally, our procedures are applied to two non-trivial empirical examples: an airline entry game and a model of trade flows.

stat.ME