arXiv · 2411.01864
Debiased Machine Learning with Many Cross-Fitting Folds
Abstract
This paper studies debiased machine learning (DML) when the number of cross-fitting folds, $K_n$, may grow with the sample size $n$. Existing fixed-$K$ asymptotic theory implies that DML1 and DML2, the two main DML variants, are asymptotically equivalent, providing no guidance on which variant to use or how to choose $K_n$. We show that this equivalence can break down when $K_n$ grows proportionally to $\sqrt{n}$: DML1 can exhibit asymptotic bias, in which case standard inference based on DML1 fails---as can occur, for instance, for the local average treatment effect (LATE)---whereas inference based on DML2 remains valid. Moreover, we show that, under an algorithmic-stability condition, estimation and inference based on DML2 are valid for any $2\le K_n \le n$, including the leave-one-out case, $K_n=n$. Finally, for scalar DML2 estimators whose first-step estimators admit a stochastic linear expansion, we derive a second-order approximation showing that larger values of $K_n$ reduce the second-order asymptotic bias and mean-squared error, although the marginal improvements diminish.
Explore related subjects
Keep this discovery
Amilcar Velez. 2024-11-04. Debiased Machine Learning with Many Cross-Fitting Folds. https://arxiv.org/abs/2411.01864
Cite the original work for its findings. Save a collection to share your selection of sources.