Searcharxiv⌕ Search

arXiv · 2610.03469

Goodness-of-Fit Testing for Groupwise Spherical Error Structures

Abstract

The analysis of large data panels is important in econometrics and beyond. Prediction and inference methods for such data typically rely on simplifying model assumptions for the covariance structure of errors. One convenient assumption is what we call groupwise sphericity: that errors are uncorrelated across individuals and have constant variance within certain groups. While theoretically useful, in large panels groupwise sphericity is often too restrictive to apply in practice. We therefore develop new quantitative inference tools to test whether deviations from this model assumption are practically relevant. Our approach covers both large-dimensional data matrices and regression panels, in a regime where the cross-sectional dimension is proportional to the sample size. The theory is based on the analysis of extreme eigenvalues of the empirical covariance matrix and uses recent advances in random matrix theory. Numerical experiments demonstrate accurate size control and good power in finite samples.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Daria Tieplova, Nina Dörnemann, Tim Kutta. 2026-10-02. Goodness-of-Fit Testing for Groupwise Spherical Error Structures. https://arxiv.org/abs/2610.03469

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

KL Convergence Guarantees for Score diffusion models under minimal data assumptions

Diffusion models are a new class of generative models that revolve around the estimation of the score function associated with a stochastic differential equation. Subsequent to its acquisition, the approximated score function is then harnessed to simulate the corresponding time-reversal process, ultimately enabling the generation of approximate data samples. Despite their evident practical significance these models carry, a notable challenge persists in the form of a lack of comprehensive quantitative results, especially in scenarios involving non-regular scores and estimators. In almost all reported bounds in Kullback Leibler (KL) divergence, it is assumed that either the score function or its approximation is Lipschitz uniformly in time. However, this condition is very restrictive in practice or appears to be difficult to establish. To circumvent this issue, previous works mainly focused on establishing convergence bounds in KL for an early stopped version of the diffusion model and a smoothed version of the data distribution, or assuming that the data distribution is supported on a compact manifold. These explorations have led to interesting bounds in either Wasserstein or Fortet-Mourier metrics. However, the question remains about the relevance of such early-stopping procedure or compactness conditions. In particular, if there exist a natural and mild condition ensuring explicit and sharp convergence bounds in KL. In this article, we tackle the aforementioned limitations by focusing on score diffusion models with fixed step size stemming from the Ornstein-Uhlenbeck semigroup and its kinetic counterpart. Our study provides a rigorous analysis, yielding simple, improved and sharp convergence bounds in KL applicable to any data distribution with finite Fisher information with respect to the standard Gaussian distribution.

math.ST↗

Estimation of conditional inequality curves and measures via estimating the conditional quantile function

In the paper conditional inequality curves and measures are proposed which allow us to describe the inequality/concentration of the conditional distribution of the feature we are interested in with respect to certain continuous variables. Moreover, for a graphical illustration of the change in values of the proposed conditional indices, a curve of conditional inequality measures is introduced. To estimate the curves and measures, a new method is proposed to estimate the conditional quantile function. This method uses quantile regression estimates for a given set of quantile orders, followed by isotonic regression on the estimated regression coefficients to ensure that the estimated conditional quantile function is nondecreasing. The consistency of the proposed estimators is proved while their finite sample performance is evaluated through simulation studies and compared with existing approaches. Finally, practical application of conditional curves and measures is demonstrated by determining estimated curves of conditional salary inequalities with respect to years of experience in different employee tenure groups, based on some real data. The code used to prepare the simulation results presented in this paper is available in a dedicated GitHub repository.

math.ST↗

Tracy-Widom, Gaussian, and Bootstrap: Approximations for Leading Eigenvalues in High-Dimensional PCA

Under certain conditions, the largest eigenvalue of a sample covariance matrix undergoes a well-known phase transition when the sample size $n$ and data dimension $p$ diverge proportionally. In the subcritical regime, this eigenvalue has fluctuations of order $n^{-2/3}$ that can be approximated by a Tracy-Widom distribution, while in the supercritical regime, it has fluctuations of order $n^{-1/2}$ that can be approximated with a Gaussian distribution. However, the statistical problem of determining which regime underlies a given dataset is far from resolved. We develop a new testing framework and procedure to address this problem. In particular, we demonstrate that the procedure has an asymptotically controlled level, and that it is power consistent for certain alternatives. Also, this testing procedure enables the design a new bootstrap method for approximating the distributions of functionals of the leading sample eigenvalues within the subcritical regime -- which is the first such method that is supported by theoretical guarantees.

math.ST↗