Searcharxiv⌕ Search

arXiv · 2610.05006

Minimax estimation of the expected conditional covariance under bounds on the covariate density

Abstract

In many scientific studies, evaluating the association between two responses requires adjusting for shared covariates. The expected conditional covariance measures this adjusted association and is used in conditional independence testing and treatment effect estimation. Researchers often estimate it using fitted conditional means, but removing the bias caused by the estimation errors of these conditional means remains a difficult task. This difficulty can increase when the true covariate density is unknown and restricted only by known upper and lower bounds. In this article, we study how accurately the expected conditional covariance can be estimated under these density bounds. We propose a multiresolution estimator that reduces bias without estimating the density. Theoretical results show that, for bounded responses and rough regression functions, an unknown density slows the optimal worst-case rate by a power of sample size compared with a known density. We establish that the optimal rate under density bounds alone strictly improves upon the previously conjectured polynomial rate by a sub-polynomial factor, resolving a long-standing conjecture. Our estimator attains this rate up to a logarithmic factor. Through simulations and a real data resampling experiment on diamonds, we show that our method can yield smaller mean squared errors than standard uncorrected estimators.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Seyoung Park. 2026-10-04. Minimax estimation of the expected conditional covariance under bounds on the covariate density. https://arxiv.org/abs/2610.05006

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Construction of optimal tests for symmetry on the torus and their quantitative error bounds

In this paper, we investigate the general problem of assessing symmetry in data points on the hyper-dimensional torus, a question that originally emerged in applications from bioinformatics and directional statistics. We develop optimal tests for symmetry for both scenarios where the center of symmetry is known and where it is unknown. Our new tests are not only valid under a given parametric hypothesis but also under a very broad class of symmetric distributions. The asymptotic behavior of the proposed tests is studied both under the null hypothesis and local alternatives. A key contribution of our paper is that we accompany our asymptotic results with error guarantees by deriving quantitative bounds on the distributional distance between the exact (unknown) distribution of the test statistic and its asymptotic counterpart by leveraging Stein's method. The finite-sample performance of the tests is evaluated through simulation studies, and their practical utility in bioinformatics is demonstrated via an application to protein folding data.

math.ST↗

Axioms for testing with data-dependent levels, e-values and p-values

The emerging literature on hypothesis testing with data-dependent and post-hoc significance levels relies on a particular extension of the Type-I error to data-dependent levels. Existing arguments for this extension are heuristic, and primarily motivated by a resulting connection to the e-value. Our first contribution is to show that it is uniquely characterized by three axioms: law-invariance, calibration to classical testing, and a mixing axiom. Inspired by a combination of Birnbaum's conditionality principle and Savage's sure-thing principle, the mixing axiom assumes that a test produced by randomly selecting between (in)valid tests must be (in)valid. Our second contribution is to show that three analogous axioms characterize the e-value as a continuous generalization of a test in a decision-theoretic framework. We recover the p-value by dropping part of the mixing axiom, showing that e-values correspond to those p-values for which a random choice between two invalid p-values cannot lead to a valid p-value. Finally, we show that the relationship between e-values and post-hoc testing goes through under much weaker axioms.

math.ST↗

Stochastic Inversion of Multivariate Uniform-Distribution-Preserving Transformations

A multivariate transformation of the unit cube with component transformations that are piecewise continuously differentiable and uniform distribution preserving (udp) is considered. A stochastic inverse transformation is defined using randomization to overcome the non-injective nature of the udp transformations. The inverse transformation preserves the uniform margins of a random vector distributed according to a copula and yields different copulas for different randomizations. A copula density transformation result for the multivariate stochastic inverse is proved and illustrated in the bivariate case.

math.ST↗