Searcharxiv⌕ Search

arXiv · 2609.39905

Is the regression $F$-test doubly robust?

Abstract

We study the robustness of the $F$-test in random design linear models, and reach a somewhat nuanced conclusion. On the positive side, one of our main results is that the size of the test is close to its nominal level as soon as either the distribution of the normalised error vector is close to uniform on the unit sphere, or the design matrix, after applying a suitable column space-preserving orthogonalisation scheme, is close to being uniformly distributed. This provides a sense in which the $F$-test is doubly robust. Our conclusion is reached by establishing a Kolmogorov to Wasserstein distance Hölder continuity property controlling the departure of the $F$-statistic from its notional $F$-distribution under the null. Writing $n$, $p$ and $p_0$ for the sample size and the dimensions of the full and null models respectively, we prove that the Hölder exponent is $1/3$ when $\min(p-p_0,n-p) = 1$ and $1/2$ when $\min(p-p_0,n-p) \geq 2$. On the other hand, these exponents are relatively small and cannot be improved in general, revealing that the size of the test may depart from its nominal level quite quickly as we move away from settings where the test is exact. In some cases, our conclusions may be improved by working with a local Kolmogorov distance that focuses on discrepancies between distribution functions in the right tail.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lucy Xia, Oliver Y. Feng, Yang Feng, Min Xu, Richard J. Samworth. 2026-09-30. Is the regression $F$-test doubly robust?. https://arxiv.org/abs/2609.39905

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Optimal Estimation of Large-Dimensional Nonlinear Factor Models

This paper studies optimal estimation of large-dimensional nonlinear factor models. The key challenge is that the observed variables are possibly nonlinear functions of some latent variables, with the functional forms left unspecified. A local principal component analysis method combining K-nearest neighbors matching and principal component analysis is proposed to estimate the factor structure and recover information on latent variables and latent functions. Large-sample properties are established, including a sharp bound on the matching discrepancy of nearest neighbors, sup-norm error bounds for estimated local factors and factor loadings, and the uniform convergence rate of the factor structure estimator. Under mild conditions our estimator of the latent factor structure can achieve the optimal rate of uniform convergence for nonparametric regression. The method is illustrated with a Monte Carlo experiment and an empirical application studying the effect of tax cuts on economic growth.

math.ST↗

Adaptive Density Estimation Using Projection Kernels and Penalized Comparison to Overfitting

In this work, we study wavelet projection estimators for density estimation, based on compactly supported $\mathcal S$-regular scaling functions. The main issue is the choice of the resolution level, which determines the bias--variance trade-off. We select this level by a Penalized Comparison to Overfitting (PCO) criterion: each candidate estimator is compared in $Ł^2(\R)$ with a single overfitting reference, and the additional stochastic fluctuation is corrected by an explicit penalty. For the selected estimator, we prove a high-probability oracle inequality and a risk oracle inequality in expectation. Over Besov balls of densities satisfying a common $Ł^\infty$ bound, the risk inequality yields the rate $n^{-2r/(2r+1)}$ uniformly over the class, with a sufficient penalty threshold that can be chosen uniformly over the class. This rate has the classical minimax order, and the procedure adapts to the unknown Besov regularity. Numerical experiments on several density shapes confirm that the data-driven level is close to the oracle one and adapts to the structure of the target.

math.ST↗

Minimax Optimal Estimation of Mean and Covariance Functions with Spectral Regularization

Estimation of the mean and covariance functions is a fundamental problem in functional data analysis, particularly for discretely observed functional data. In this work, we study a regularization-based framework for estimating the mean and the covariance functions within a reproducing kernel Hilbert space (RKHS) setting. Our approach utilizes a spectral regularization technique under Hölder-type source conditions, allowing for a broad class of regularization schemes and accommodating a wide range of smoothness assumptions on the target functions. In contrast to RKHS formulations that assume the target belongs to the underlying RKHS, our source-condition framework also accommodates misspecified targets. Convergence rates for the proposed estimators are derived, and we derive corresponding minimax lower bounds and identify regimes in which the upper bounds are optimal, or optimal up to logarithmic factors.

math.ST↗