Searcharxiv⌕ Search

arXiv · 2609.37412

Debiased Inference for Bounding Wage Inequality with Many Controls

Abstract

We study estimation and inference for a partially identified parameter whose identified set depends on a first-stage nuisance parameter that must itself be estimated. Combining the criterion-function approach with the theory of Neyman-orthogonal moments that underlies double/debiased machine learning, we propose a two-step procedure: the point-identified nuisance is estimated by flexible machine-learning methods, and the set-identified target is recovered as a level set of a sample criterion built from orthogonal moment inequalities with cross-fitting. When the contour level is bounded, we show that the resulting set estimator converges in Hausdorff distance at the parametric rate of the infeasible criterion built on the true nuisance. We further develop a subsampling procedure that delivers asymptotically valid coverage, provided the product of the first-stage estimation errors is $o(N^{-1/2})$. We illustrate the method on bounds for the wage distribution and the interquantile range under selection into employment and on the gender wage gap with an interval-censored wage. The empirical application studies the gender wage gap using the March supplement of the 2015 Current Population Survey.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yaroslav Korobka, Vira Semenova. 2026-09-29. Debiased Inference for Bounding Wage Inequality with Many Controls. https://arxiv.org/abs/2609.37412

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

When Is Neyman-Orthogonal Inference Feasible? Existence and Relevance

Neyman-orthogonal moments underpin debiased and double machine learning, yet their existence and informativeness are typically presumed. We ask three questions: when do nontrivial orthogonal moments exist, which target directions contain first-order information, and does a proposed moment detect all such directions? Restricted local overidentification characterizes existence under attainability. Effective Fisher information determines relevance; its rank counts informative target directions. A rank sandwich combines moment Jacobians with target changes offset by nuisance perturbations, distinguishing information missed by a moment from information absent from the experiment. In sample selection without covariate exclusions, this yields identification-robust inference under rank deficiency. In latent-variable models, the informative target space is an RKHS. For Gaussian repeated measurements, additional measurements enlarge the class of smoothed distributional targets that are regularly learnable, while CDFs, quantiles, and threshold-policy functionals remain uninformative at the parametric rate for every finite number of measurements.

econ.EM↗

A Dimension-Agnostic Bootstrap Anderson-Rubin Test For Instrumental Variable Regressions

Weak-identification-robust tests for instrumental variable (IV) regressions are typically developed separately depending on whether the number of IVs is treated as fixed or increasing with the sample size, forcing researchers to make a stance on the asymptotic behavior, which is often ambiguous in practice. This paper proposes a bootstrap-based, dimension-agnostic Anderson-Rubin (AR) test that achieves correct asymptotic size regardless of whether the number of IVs is fixed or diverging, and even accommodates cases where the number of IVs exceeds the sample size. By incorporating ridge regularization, our approach reduces the effective rank of the projection matrix and yields regimes where the limiting distribution of the AR statistic can be a weighted chi-squared, a normal, or a mixture of the two. Strong approximation results ensure that the bootstrap procedure remains uniformly valid across all regimes, while also delivering substantial power gains over existing methods by exploiting rank reduction.

econ.EM↗

Debiased Machine Learning for Unobserved Heterogeneity: High-Dimensional Panels and Measurement Error Models

Developing robust inference for models with nonparametric Unobserved Heterogeneity (UH) is both important and challenging. We propose novel Debiased Machine Learning (DML) procedures for valid inference on functionals of UH, allowing for partial identification of multivariate target and high-dimensional nuisance parameters. Our main contribution is a full characterization of all relevant Neyman-orthogonal moments in models with nonparametric UH, where relevance means informativeness about the parameter of interest. Under additional support conditions, orthogonal moments are globally robust to the distribution of the UH. They may still involve other high-dimensional nuisance parameters, but their local robustness reduces regularization bias and enables valid DML inference. We apply these results to: (i) common parameters, average marginal effects, and variances of UH in panel data models with high-dimensional controls; (ii) moments of the common factor in the Kotlarski model with a factor loading; and (iii) smooth functionals of teacher value-added. Monte Carlo simulations show substantial efficiency gains from using efficient orthogonal moments relative to ad-hoc choices. We illustrate the practical value of our approach by showing that existing estimates of the average and variance effects of maternal smoking on child birth weight are robust.

econ.EM↗