SearcharxivSearch

arXiv subjects

Victor Chernozhukov

Publications and source records attributed to Victor Chernozhukov.

At least 19 recordsLinked to original sources

Conditional Rank-Rank Regression

Rank-rank regression is commonly employed in economic research as a way of capturing the relationship between two economic variables. The slope of this regression is the Spearman rank correlation, a classical measure of association. However, in many applications it is common practice to include covariates to account for differences in association levels between groups as defined by the values of these covariates. This is either done by including the covariates or by modeling the residuals obtained after partialing out the impact of the covariates. In each of these instances the resulting rank-rank regression coefficients can be difficult to interpret. We propose the conditional rank-rank regression, which uses conditional ranks instead of unconditional ranks, to measure average within-group persistence. The coefficient of this new regression corresponds to the average Spearman rank correlation conditional on the covariates, a natural summary measure of within-group association. We develop a flexible estimation approach using distribution regression and establish a theoretical framework for large sample inference. An empirical study on intergenerational income mobility in Switzerland demonstrates the advantages of this approach. The study reveals stronger intergenerational persistence between fathers and sons compared to fathers and daughters, with the within-group persistence being 62% of the overall income persistence for sons and 52% for daughters. Smaller families and those with highly educated fathers exhibit greater persistence in economic status.

econ.EM

Omitted Variable Bias in Difference-in-Differences Designs

We study the omitted variable bias (OVB) problem in canonical difference-in-differences (DiD) designs when unobserved confounding induces departures from the parallel trends assumption. Our results provide a novel characterization of the OVB formula for the average treatment effect on the treated (ATT), which is of independent interest. We show how the ATT bias is mainly governed by the strength of confounding in the treatment assignment mechanism and provide alternative ways of quantifying this strength, such as (i) changes in the average odds of treatment among the treated, (ii) confounding imbalance between treated and control units, or (iii) variation explained in treatment odds among the untreated. Building on these results, we offer sensitivity statistics for routine reporting, describing the minimum strength of confounding required to overturn the conclusions of a DiD study, as well as formal bounds on the strength of confounders based on comparisons to observed covariates or pre-trends. Finally, we provide flexible and efficient statistical inference methods for the bounds on ATT, which can leverage modern machine learning algorithms for estimation. We demonstrate the utility of our approach in an empirical example that estimates the effects of minimum wage on teen employment.

stat.ME

Estimating Causal Effects of Discrete and Continuous Treatments with Binary Instruments

We propose an instrumental variable framework for identifying and estimating causal effects of discrete and continuous treatments with binary instruments. The basis of our approach is a local copula representation of the joint distribution of the potential outcomes and unobservables determining treatment assignment. This representation allows us to introduce an identifying assumption, so-called constant local dependence, that restricts the local dependence of the copula with respect to the treatment propensity. We show that constant local dependence identifies treatment effects for the entire population and other subpopulations such as the treated. The identification results are constructive and lead to practical estimation and inference procedures based on distribution regression. An application to estimating the effect of sleep on well-being uncovers interesting patterns of heterogeneity.

econ.EM

Linear Estimation of Structural and Causal Effects for Nonseparable Panel Data

This paper develops linear estimators for structural and causal parameters of nonseparable models using panel data. These models incorporate unobserved, time-varying, individual heterogeneity, which may be correlated with the regressors. Estimation is based on an approximation of a conditional average potential outcome by a linear sieve specification with individual-specific parameters. Effects of interest are estimated by a bias corrected average of individual ridge regressions. We demonstrate how this approach can be applied to estimate causal effects, counterfactual consumer welfare, and averages of individual taxable income elasticities. We show that the proposed estimator has an empirical Bayes interpretation and possesses a number of other useful properties. We formulate Large-$T$ asymptotics that can accommodate discrete regressors and which bypass partial identification in this case. We employ the methods to estimate average equivalent variation and deadweight loss for potential price increases using data on grocery purchases.

econ.EM

Arellano-Bond LASSO Estimator for Dynamic Linear Panel Models

The Arellano-Bond estimator is a fundamental method for dynamic panel data models, widely used in practice. It can be severely biased when the time series dimension of the data, $T$, is long. The source of the bias is the large degree of overidentification. We propose a simple two-step approach to deal with this problem. The first step applies LASSO to the cross-section data at each time period to select the most informative moment conditions, exploiting the approximately sparse structure of these conditions. The second step applies a linear instrumental variable estimator using the instruments constructed from the moment conditions selected in the first step. Using asymptotic sequences where the two dimensions of the panel grow with the sample size, we show that the new estimator is consistent and asymptotically normal under much weaker conditions on $T$ than the Arellano-Bond estimator. Our theory covers models with high-dimensional covariates including multiple lags of the dependent variable and strictly exogenous covariates, which are becoming common in modern applications. We illustrate our approach by applying it to weekly county-level panel data from the United States to study opening K-12 schools and other mitigation policies' short and long-term effects on COVID-19's spread.

econ.EM

Automatic Debiased Machine Learning for Dynamic Treatment Effects and General Nested Functionals

Many canonical models in causal inference and structural econometrics have recursive identification formulas. In causal inference, recursion arises when identification requires both pre- and post-treatment covariates. For example, short-term surrogate outcomes are measured after the treatment, and serve as necessary covariates when identifying long-term effects. Post-treatment covariates are also required for identification of dynamic difference-in-differences designs, time-varying treatment regimes, and mediation analysis. In structural econometrics, recursion arises through evolving state variables, for example in dynamic sample selection models and dynamic discrete choice models. In this paper, we propose an automatic and recursive method for inference, applicable to such formulas, allowing for flexible estimation by neural networks and random forests. As a technical contribution, we introduce recursive Riesz representers.

econ.EM

Agentic Economic Modeling

We introduce Agentic Economic Modeling (AEM), a framework that aligns synthetic LLM choices with small-sample human evidence for econometric inference. AEM first generates task-conditioned synthetic choices via LLMs, then learns a bias-correction mapping from task features and raw LLM choices to human-aligned choices, upon which standard econometric estimators perform inference to recover demand elasticities and treatment effects. We validate AEM in two experiments. In a large scale conjoint study, using only 10% of the original data to fit the correction model lowers the error of the demand-parameter estimates, while uncorrected LLM choices increase the errors. In a regional field experiment, a mixture model calibrated on 10% of geographic regions estimates a treatment effect of -65$\pm$10 bps on the hold-out regions, closely matching the full human experiment (-60$\pm$8 bps). These results demonstrate AEM's potential to improve RCT efficiency and represent a step toward LLM-based counterfactual generation.

econ.EM

Plausible GMM: A Quasi-Bayesian Approach

Structural estimation in economics often makes use of models formulated in terms of moment conditions. While these moment conditions are generally well-motivated, it is often unknown whether the moment restrictions hold exactly. We consider a framework where researchers model their belief about the potential degree of misspecification via a prior distribution and adopt a quasi-Bayesian approach for performing inference on structural parameters. We provide quasi-posterior concentration results, verify that quasi-posteriors can be used to obtain approximately optimal Bayesian decision rules under the maintained prior structure over misspecification, and provide a form of frequentist coverage results. We illustrate the approach through empirical examples where we obtain informative inference for structural objects allowing for substantial relaxations of the requirement that moment conditions hold exactly.

econ.EM

An Introduction to Double/Debiased Machine Learning

This paper provides an introduction to Double/Debiased Machine Learning (DML). DML is a general approach to performing inference about a target parameter in the presence of nuisance functions: objects that are needed to identify the target parameter but are not of primary interest. Nuisance functions arise naturally in many settings, such as when controlling for confounding variables or leveraging instruments. The paper describes two biases that arise from nuisance function estimation and explains how DML alleviates these biases. Consequently, DML allows the use of flexible methods, including machine learning tools, for estimating nuisance functions, reducing the dependence on auxiliary functional form assumptions and enabling the use of complex non-tabular data, such as text or images. We illustrate the application of DML through simulations and empirical examples. We conclude with a discussion of recommended practices. A companion website includes additional examples with code and references to other resources.

econ.EM

Adventures in Demand Analysis Using AI

This paper advances empirical demand analysis by integrating multimodal product representations derived from artificial intelligence (AI). Using a detailed dataset of toy cars on textit{Amazon.com}, we combine text descriptions, images, and tabular covariates to represent each product using transformer-based embedding models. These embeddings capture nuanced attributes, such as quality, branding, and visual characteristics, that traditional methods often struggle to summarize. Moreover, we fine-tune these embeddings for causal inference tasks. We show that the resulting embeddings substantially improve the predictive accuracy of sales ranks and prices and that they lead to more credible causal estimates of price elasticity. Notably, we uncover strong heterogeneity in price elasticity driven by these product-specific features. Our findings illustrate that AI-driven representations can enrich and modernize empirical demand analysis. The insights generated may also prove valuable for applied causal inference more broadly.

econ.GN

Hedonic Prices and Quality Adjusted Price Indices Powered by AI

We develop empirical models that efficiently process large amounts of unstructured product data (text, images, prices, quantities) to produce accurate hedonic price estimates and derived indices. To achieve this, we generate abstract product attributes (or ``features'') from descriptions and images using deep neural networks. These attributes are then used to estimate the hedonic price function. To demonstrate the effectiveness of this approach, we apply the models to Amazon's data for first-party apparel sales, and estimate hedonic prices. The resulting models have a very high out-of-sample predictive accuracy, with $R^2$ ranging from $80\%$ to $90\%$. Finally, we construct the AI-based hedonic Fisher price index, chained at the year-over-year frequency, and contrast it with the CPI and other electronic indices.

econ.GN

Policy Learning with Confidence

This paper introduces a rule for policy selection in the presence of estimation uncertainty, explicitly accounting for estimation risk. The rule belongs to the class of risk-aware rules on the efficient decision frontier, characterized as policies offering maximal estimated welfare for a given level of estimation risk. Among this class, the proposed rule is chosen to provide a reporting guarantee, ensuring that the welfare delivered exceeds a threshold with a pre-specified confidence level. We apply this approach to the allocation of a limited budget among social programs using estimates of their marginal value of public funds and associated standard errors.

econ.EM

Automatic debiased machine learning and sensitivity analysis for sample selection models

In this paper, we extend the Riesz representation framework to causal inference under sample selection, where both treatment assignment and outcome observability are non-random. Formulating the problem in terms of a Riesz representer enables stable estimation and a transparent decomposition of omitted variable bias into three interpretable components: a data-identified scale factor, outcome confounding strength, and selection confounding strength. For estimation, we employ the ForestRiesz estimator, which accounts for selective outcome observability while avoiding the instability associated with direct propensity score inversion. We assess finite-sample performance through a simulation study and show that conventional double machine learning approaches can be highly sensitive to tuning parameters due to their reliance on inverse probability weighting, whereas the ForestRiesz estimator delivers more stable performance by leveraging automatic debiased machine learning. In an empirical application to the gender wage gap in the U.S., we find that our ForestRiesz approach yields larger treatment effect estimates than a standard double machine learning approach, suggesting that ignoring sample selection leads to an underestimation of the gender wage gap. Sensitivity analysis indicates that implausibly strong unobserved confounding would be required to overturn our results. Overall, our approach provides a unified, robust, and computationally attractive framework for causal inference under sample selection.

econ.EM

Sensitivity Analysis for Causal ML: A Use Case at Booking.com

Causal Machine Learning has emerged as a powerful tool for flexibly estimating causal effects from observational data in both industry and academia. However, causal inference from observational data relies on untestable assumptions about the data-generating process, such as the absence of unobserved confounders. When these assumptions are violated, causal effect estimates may become biased, undermining the validity of research findings. In these contexts, sensitivity analysis plays a crucial role, by enabling data scientists to assess the robustness of their findings to plausible violations of unconfoundedness. This paper introduces sensitivity analysis and demonstrates its practical relevance through a (simulated) data example based on a use case at Booking.com. We focus our presentation on a recently proposed method by Chernozhukov et al. (2023), which derives general non-parametric bounds on biases due to omitted variables, and is fully compatible with (though not limited to) modern inferential tools of Causal Machine Learning. By presenting this use case, we aim to raise awareness of sensitivity analysis and highlight its importance in real-world scenarios.

econ.EM

Automatic Debiased Machine Learning for Covariate Shifts

We present machine learning estimators for causal and predictive parameters under covariate shift, where covariate distributions differ between training and target populations. One such parameter is the average effect of a policy that alters the covariate distribution, such as a treatment modifying surrogate covariates used to predict long-term outcomes. Another example is the average treatment effect for a population with a shifted covariate distribution, like the effect of a policy on the treated group. We propose a debiased machine learning method to estimate a broad class of these parameters in a statistically reliable and automatic manner. Our method eliminates regularization biases arising from the use of machine learning tools in high-dimensional settings, relying solely on the parameter's defining formula. It employs data fusion by combining samples from target and training data to eliminate biases. We prove that our estimator is consistent and asymptotically normal. Computational experiments and an empirical study on the impact of minimum wage increases on teen employment--using the difference-in-differences framework with unconfoundedness--demonstrate the effectiveness of our method.

stat.ME

Bivariate Distribution Regression; Theory, Estimation and an Application to Intergenerational Mobility

We employ distribution regression (DR) to estimate the joint distribution of two outcome variables conditional on chosen covariates. While Bivariate Distribution Regression (BDR) is useful in a variety of settings, it is particularly valuable when some dependence between the outcomes persists after accounting for the impact of the covariates. Our analysis relies on a result from Chernozhukov et al. (2018) which shows that any conditional joint distribution has a local Gaussian representation. We describe how BDR can be implemented and present some associated functionals of interest. As modeling the unexplained dependence is a key feature of BDR, we focus on functionals related to this dependence. We decompose the difference between the joint distributions for different groups into composition, marginal and sorting effects. We provide a similar decomposition for the transition matrices which describe how location in the distribution in one of the outcomes is associated with location in the other. Our theoretical contributions are the derivation of the properties of these estimated functionals and appropriate procedures for inference. Our empirical illustration focuses on intergenerational mobility. Using the Panel Survey of Income Dynamics data, we model the joint distribution of parents' and children's earnings. By comparing the observed distribution with constructed counterfactuals, we isolate the impact of observable and unobservable factors on the observed joint distribution. We also evaluate the forces responsible for the difference between the transition matrices of sons' and daughters'.

econ.EM

Minimax Semiparametric Learning With Approximate Sparsity

Estimating linear, mean-square continuous functionals is a pivotal challenge in statistics. In high-dimensional contexts, this estimation is often performed under the assumption of exact model sparsity, meaning that only a small number of parameters are precisely non-zero. This excludes models where linear formulations only approximate the underlying data distribution, such as nonparametric regression methods that use basis expansion such as splines, kernel methods or polynomial regressions. Many recent methods for root-$n$ estimation have been proposed, but the implications of exact model sparsity remain largely unexplored. In particular, minimax optimality for models that are not exactly sparse has not yet been developed. This paper formalizes the concept of approximate sparsity through classical semi-parametric theory. We derive minimax rates under this formulation for a regression slope and an average derivative, finding these bounds to be substantially larger than those in low-dimensional, semi-parametric settings. We identify several new phenomena. We discover new regimes where rate double robustness does not hold, yet root-$n$ estimation is still possible. In these settings, we propose an estimator that achieves minimax optimal rates. Our findings further reveal distinct optimality boundaries for ordered versus unordered nonparametric regression estimation.

math.ST

Uniform Inference on High-dimensional Spatial Panel Networks

We propose employing a high-dimensional generalized method of moments (GMM) estimator, regularized for dimension reduction and subsequently debiased to correct for shrinkage bias (referred to as a debiased-regularized estimator), for inference on large-scale spatial panel networks. In particular, the network structure, which incorporates a flexible sparse deviation that can be regarded either as a latent component or as a misspecification of a predetermined adjacency matrix, is estimated using a debiased machine learning approach. The theoretical analysis establishes the consistency and asymptotic normality of our proposed estimator, taking into account general temporal and spatial dependencies inherent in the data-generating processes. A primary contribution of our study is the development of a uniform inference theory, which enables hypothesis testing on the parameters of interest, including zero or non-zero elements in the network structure. Additionally, the asymptotic properties of the estimator are derived for both linear and nonlinear moments. Simulations demonstrate the superior performance of our proposed approach. Finally, we apply our methodology to investigate the spatial network effects of stock returns.

econ.EM