Searcharxiv⌕ Search

arXiv · 2610.02944

Cross-Fitting Under Nonregularity: Normality and Inference via Locality

Abstract

Cross-fitting is routine in much of applied research. While conventional confidence intervals that ignore cross-fold dependence are asymptotically valid in several settings, they undercover in many applications that share a common form of nonregularity: from the classic cross-validation problem of testing whether a fitted model outperforms another, to testing for heterogeneous treatment effects with machine learning, to estimating the value of a potentially non-unique optimal treatment regime. Exploiting a new locality condition, I show that a large class of cross-fitting estimators still satisfies a central limit theorem despite the nonregularity, but with an asymptotic variance that must be adjusted for the cross-fold correlation. Then, I propose a method for estimating this correlation and construct new confidence intervals that attain asymptotically nominal coverage. Finally, I show that the proposed confidence intervals attain approximately nominal coverage in a simulation study with random forests and neural networks.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bruno Fava. 2026-10-02. Cross-Fitting Under Nonregularity: Normality and Inference via Locality. https://arxiv.org/abs/2610.02944

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Endogenous Control Bias and Correction with Nonparametric Methods

Controls are often needed to make a treatment or instrument conditionally valid, yet they may be related to unobserved outcome determinants. We show that restricted parametric adjustment can fail to recover the structural average response even when the structural model is correctly specified. Under conditional independence and sufficient conditional treatment variation, flexible conditional mean methods identify the structural average response in a nonseparable model, while a control function extension accommodates endogenous treatments with conditionally valid instruments. We develop a specification test to guide the choice between restricted and flexible functional forms. Simulations and a China shock application illustrate the methods.

econ.EM↗

Interpretable Discovery from Unstructured Data: A High-Dimensional Approach

We propose an automatic, general-purpose framework for making discoveries from unstructured data (e.g., text data from open-ended surveys of economic beliefs). The framework leverages recent methods from the literature on AI interpretability to transform unstructured datasets into high-dimensional, structured datasets of interpretable concept measurements; specifies concept-level parameters and null hypotheses based on this transformed dataset; tests these hypotheses using algorithms validated by new results in high-dimensional multiple testing, producing a selected set ("discoveries"); and both generates and evaluates human-interpretable natural language descriptions of these discoveries. The proposed framework has few researcher degrees of freedom, is robust to data snooping, mitigates under-exploration, and facilitates fast and inexpensive sensitivity analysis and replication. We revisit applications to recent descriptive and causal analyses of unstructured data in empirical economics, and find this framework is able to automatically replicate existing discoveries, add nuance to others, and make entirely new discoveries as well.

econ.EM↗

Pragmatic DML with AI-Learned Representations

Text, images, and other rich covariates are increasingly compressed into AI-learned representations and then used as controls in causal analysis. We study when this approach is valid and develop a practical framework for causal inference with learned representations. For a broad class of estimands, an imperfect representation distorts the target causal parameter by the product of two representation errors: one in the outcome regression and one in the balancing weight (or Riesz representer). This yields three constructive results. First, cross-fitted double machine learning (DML) provides valid Wald inference for the representation-dependent target. When representation errors are small, the same interval covers the causal parameter, and it can even attain the semiparametric efficiency bound. Second, fold-wise representation learning (or fine-tuning) is compatible with DML inference for the causal parameter. To this end, we develop convex- and star-aggregation pipelines for learning and combining representations. Third, when representation errors are substantial, we can provide interpretable sensitivity regions and root-$n$ inference for their endpoints. In a multi-modal demand application, seven representation-specific estimates and their star aggregate all imply a negative near-unit elasticity for rank-based price response, and the result remains robust over the reported sensitivity grid.

econ.EM↗