SearcharxivSearch

arXiv subjects

Matey Neykov

Publications and source records attributed to Matey Neykov.

At least 19 recordsLinked to original sources

Sparse Convexification for High-Dimensional Constrained Regression

We study high-dimensional linear regression under a general symmetric convex constraint. Rather than imposing a specific sparsity-inducing penalty, we start from an arbitrary sign-symmetric and permutation-invariant convex body $K\subseteq \mathbb R^p$ and construct the sparse convexification hierarchy \[ K^{(s)} = \operatorname{conv}\{v\in K:\|v\|_0\le s\}. \] We propose a penalized least-squares estimator that searches over this hierarchy and adapts to the best sparse convex approximation of the target. Under standard sub-Gaussian assumptions on the random design and noise, we prove an oracle inequality showing that the estimator adapts to the best sparse convex approximation of the target. For an $s$-sparse target, the result yields a squared-error rate governed by the noise level $\sigma$, and the Gaussian width of the sparse convexification $K^{(s)}$. The method applies broadly to symmetric norm balls and can be implemented using oracle access to the Minkowski functional of $K$. As a special case, the framework yields a consistency result for the constrained Lasso.

math.ST

Fast Near-Optimal Estimation over Symmetric Norm Balls

This short note proposes a polynomial-time algorithm for near-optimal Euclidean estimation of a signal constrained to lie in the unit ball of a symmetric norm, where the symmetry is with respect to a known basis and the norm is accessible through an evaluation oracle. We further extend the method to a random-design, moderate-dimensional linear regression setting, where the regression parameter is likewise assumed to belong to a constraint set defined by a symmetric norm.

math.ST

Efficient Robust Constrained Signal Detection via Kolmogorov Width Approximations

Robust statistical inference often faces a severe computational-statistical gap when dealing with complex parameter spaces. We investigate minimax signal detection in the Gaussian sequence model under strong $\epsilon$-contamination, where the signal belongs to a general prior constraint $K$. Existing optimal tests require computing the exact Kolmogorov $k$-width of $K$, a computationally intractable task for general non-trivial sets. We bridge this gap by proposing a polynomial-time testing framework that universally applies to balanced, type-2, and exactly 2-convex constraints. By leveraging a semidefinite programming relaxation and a modified ellipsoid method equipped with an approximate subgradient oracle, we efficiently approximate the Kolmogorov widths. Remarkably, our unconditional efficient algorithm achieves a robust detection boundary that matches existing upper bounds up to a mere polylogarithmic factor. This establishes a computationally tractable testing solution for a broad class of structured signals without requiring prior knowledge of their exact geometric complexity.

math.ST

Robust mean estimation under star-shaped constraints with heavy-tailed noise

We study the problem of robust mean estimation with adversarially contaminated data under star-shaped constraints in a heavy-tailed noise setting, where only a finite second moment $ \sigma ^2 $ is assumed. For a contamination level $ \varepsilon$ below some constant, we show that the minimax rate of the squared $ \ell_2 $ loss is $ \max( \delta ^{*2}, \varepsilon \sigma ^2) \wedge d^2 $ for a star-shaped set with diameter $ d $ (set $d = \infty$ if the set is unbounded), with $ \delta ^* $ determined via the local entropy $ \log M^\mathrm{ loc }(\delta ,c) $ as \begin{align*} \delta ^*:= \sup\bigg\{\delta \geq 0: N\frac{\delta ^2}{\sigma ^2}\leq \log M^\mathrm{ loc }(\delta ,c) \bigg\}, \end{align*} where $ c $ is a sufficiently large constant. Crucially, we require that the sample size satisfies $N \gtrsim \mathop{ \sup }\limits_{\delta \geq 0} \log M^\mathrm{ loc }(\delta ,c)$. We also show that the minimax rate is $ \max(\delta^{*2},\varepsilon ^2\sigma ^2) \wedge d^2 $ for known or sign-symmetric distributions, matching the rate achieved in the Gaussian case.

math.ST

Polynomial-Time Near-Optimal Estimation over Certain Type-2 Convex Bodies

We develop polynomial-time algorithms for near-optimal minimax mean estimation under $\ell_2$-squared loss in a Gaussian sequence model under convex constraints. The parameter space is an origin-symmetric, type-2 convex body $K \subset \mathbb{R}^n$, and we assume additional regularity conditions: specifically, we assume $K$ is well-balanced, i.e., there exist known radii $r, R > 0$ such that $r B_2 \subseteq K \subseteq R B_2$, as well as oracle access to the Minkowski gauge of $K$. Under additional conditions guaranteeing an efficient approximate quadratic-form-maximization oracle on $K$, our procedures achieve the minimax rate up to factors that depend polylogarithmically on the dimension, while remaining computationally efficient. We further extend our methodology to the linear regression and robust heavy-tailed settings, establishing polynomial-time near-optimal estimators when the constraint set satisfies the regularity conditions above. To the best of our knowledge, these results provide the first general framework for attaining statistically near-optimal performance under such broad geometric constraints while preserving computational tractability.

math.ST

Minimaxity and Efficiency in Exponential Family Regression: From Star-Shaped to Convex Constraints

This paper establishes the minimax estimation rate for nonparametric exponential family regression under star-shaped constraints. We consider a parameter space $K$ that is a star-shaped subset of the hypercube $[-M, M]^n$ for a known constant $M > 0$. We operate under the assumption that the underlying exponential family is nonsingular with a twice continuously differentiable log-partition function. Our main result demonstrates that the minimax rate of the $\ell _{2}$ error of such estimation problem is $\epsilon^{*2} \wedge \operatorname{diam}(K)^2$ up to constants exclusively depending on $M$. Here, the critical radius $\epsilon^*$ is defined as \begin{equation*} \epsilon^* = \sup \{\epsilon \left\lvert\right. \epsilon^2 \kappa(M) \le \log N^{\text{loc}}(\epsilon,c)\}, \end{equation*} where $N^{\text{loc}}(\epsilon,c)$ denotes the local metric entropy of $K$, and $\kappa(M) > 0, c>0$ are constants depending only on $M$. Such minimax rate is established by a match between an information-theoretic lower bound and an upper bound implied by a theoretical algorithm. Furthermore, we investigate the computational aspects of this estimation problem. Under mildly stronger assumptions on the constraint set $K$, we propose a computationally efficient, polynomial-time algorithm. We prove that the resulting estimator achieves the minimax optimal rate up to poly-logarithmic factors in the dimension $n$ and the geometric parameters of $K$. Finally, to illustrate the efficacy of our framework, we derive the minimax optimal rates for some concrete examples.

math.ST

Robust density estimation over star-shaped density classes

We establish a novel criterion for comparing the performance of two densities, $g_1$ and $g_2$, within the context of corrupted data. Utilizing this criterion, we propose an algorithm to construct a density estimator within a star-shaped density class, $\mathcal{F}$, under conditions of data corruption. We proceed to derive the minimax upper and lower bounds for density estimation across this star-shaped density class, characterized by densities that are uniformly bounded above and below (in the sup norm), in the presence of adversarially corrupted data. Specifically, we assume that a fraction $ε\leq \frac{1}{3}$ of the $N$ observations are arbitrarily corrupted. We obtain the minimax upper bound $\max\{ τ_{\overline{J}}^2, ε\} \wedge d^2$. Under certain conditions, we obtain the minimax risk, up to proportionality constants, under the squared $L_2$ loss as $$ \max\left\{ τ^{*2} \wedge d^2, ε\wedge d^2 \right\}, $$ where $τ^* := \sup\left\{ τ: Nτ^2 \leq \log \mathcal{M}_{\mathcal{F}}^{\text{loc}}(τ, c) \right\}$ for a sufficiently large constant $c$. Here, $\mathcal{M}_{\mathcal{F}}^{\text{loc}}(τ, c)$ denotes the local entropy of the set $\mathcal{F}$, and $d$ is the $L_2$ diameter of $\mathcal{F}$.

math.ST

Information theoretic limits of robust sub-Gaussian mean estimation under star-shaped constraints

We obtain the minimax rate for a mean location model with a bounded star-shaped set $K \subseteq \mathbb{R}^n$ constraint on the mean, in an adversarially corrupted data setting with Gaussian noise. We assume an unknown fraction $\epsilon \le 1/2-\kappa$ for some fixed $\kappa\in(0,1/2]$ of $N$ observations are arbitrarily corrupted. We obtain a minimax risk up to proportionality constants under the squared $\ell_2$ loss of $\max(\eta^{*2},\sigma^2\epsilon^2)\wedge d^2$ with \begin{align*} \eta^* = \sup \bigg\{\eta \ge 0 : \frac{N\eta^2}{\sigma^2} \leq \log \mathcal{M}_K^{\operatorname{loc}}(\eta,c)\bigg\}, \end{align*} where $\log \mathcal{M}_K^{\operatorname{loc}}(\eta,c)$ denotes the local entropy of the set $K$, $d$ is the diameter of $K$, $\sigma^2$ is the variance, and $c$ is some sufficiently large absolute constant. A variant of our algorithm achieves the same rate for settings with known or symmetric sub-Gaussian noise, with a smaller breakdown point, still of constant order. We further study the case of unknown sub-Gaussian noise and show that the rate is slightly slower: $\max(\eta^{*2},\sigma^2\epsilon^2\log(1/\epsilon))\wedge d^2$. We generalize our results to the case when $K$ is star-shaped but unbounded.

math.ST

Some facts about the optimality of the LSE in the Gaussian sequence model with convex constraint

We consider a convex constrained Gaussian sequence model and characterize necessary and sufficient conditions for the least squares estimator (LSE) to be minimax optimal. For a closed convex set $K\subset \mathbb{R}^n$ we observe $Y=\mu+\xi$ for $\xi\sim \mathcal{N}(0,\sigma^2\mathbb{I}_n)$ and $\mu\in K$ and aim to estimate $\mu$. We characterize the worst case risk of the LSE in multiple ways by analyzing the behavior of the local Gaussian width on $K$. We demonstrate that optimality is equivalent to a Lipschitz property of the local Gaussian width mapping. We also provide theoretical algorithms that search for the worst case risk. We then provide examples showing optimality or suboptimality of the LSE on various sets, including $\ell_p$ balls for $p\in[1,2]$, pyramids, solids of revolution, and multivariate isotonic regression, among others.

math.ST

Semi-Supervised U-statistics

Semi-supervised datasets are ubiquitous across diverse domains where obtaining fully labeled data is costly or time-consuming. The prevalence of such datasets has consistently driven the demand for new tools and methods that exploit the potential of unlabeled data. Responding to this demand, we introduce semi-supervised U-statistics enhanced by the abundance of unlabeled data, and investigate their statistical properties. We show that the proposed approach is asymptotically Normal and exhibits notable efficiency gains over classical U-statistics by effectively integrating various powerful prediction tools into the framework. To understand the fundamental difficulty of the problem, we derive minimax lower bounds in semi-supervised settings and showcase that our procedure is semi-parametrically efficient under regularity conditions. Moreover, tailored to bivariate kernels, we propose a refined approach that outperforms the classical U-statistic across all degeneracy regimes, and demonstrate its optimality properties. Simulation studies are conducted to corroborate our findings and to further demonstrate our framework.

math.ST

Characterizing the minimax rate of nonparametric regression under bounded star-shaped constraints

We quantify the minimax rate for a nonparametric regression model over a star-shaped function class $\mathcal{F}$ with bounded diameter. We obtain a minimax rate of ${\varepsilon^{\ast}}^2\wedge\mathrm{diam}(\mathcal{F})^2$ where \[\varepsilon^{\ast} =\sup\{\varepsilon\ge 0:n\varepsilon^2 \le \log M_{\mathcal{F}}^{\operatorname{loc}}(\varepsilon,c)\},\] where $\log M_{\mathcal{F}}^{\operatorname{loc}}(\cdot, c)$ is the local metric entropy of $\mathcal{F}$, $c$ is some absolute constant scaling down the entropy radius, and our loss function is the squared population $L_2$ distance over our input space $\mathcal{X}$. In contrast to classical works on the topic [cf. Yang and Barron, 1999], our results do not require functions in $\mathcal{F}$ to be uniformly bounded in sup-norm. In fact, we propose a condition that simultaneously generalizes boundedness in sup-norm and the so-called $L$-sub-Gaussian assumption that appears in the prior literature. In addition, we prove that our estimator is adaptive to the true point in the convex-constrained case, and to the best of our knowledge this is the first such estimator in this general setting. This work builds on the Gaussian sequence framework of Neykov [2022] using a similar algorithmic scheme to achieve the minimax rate. Our algorithmic rate also applies with sub-Gaussian noise. We illustrate the utility of this theory with examples including multivariate monotone functions, linear functionals over ellipsoids, and Lipschitz classes.

math.ST

Conditional Independence Testing for Discrete Distributions: Beyond $χ^2$- and $G$-tests

This paper is concerned with the problem of conditional independence testing for discrete data. In recent years, researchers have shed new light on this fundamental problem, emphasizing finite-sample optimality. The non-asymptotic viewpoint adapted in these works has led to novel conditional independence tests that enjoy certain optimality under various regimes. Despite their attractive theoretical properties, the considered tests are not necessarily practical, relying on a Poissonization trick and unspecified constants in their critical values. In this work, we attempt to bridge the gap between theory and practice by reproving optimality without Poissonization and calibrating tests using Monte Carlo permutations. Along the way, we also prove that classical asymptotic $χ^2$- and $G$-tests are notably sub-optimal in a high-dimensional regime, which justifies the demand for new tools. Our theoretical results are complemented by experiments on both simulated and real-world datasets. Accompanying this paper is an R package UCI that implements the proposed tests.

math.ST

Revisiting Le Cam's Equation: Exact Minimax Rates over Convex Density Classes

We study the classical problem of deriving minimax rates for density estimation over convex density classes. Building on the pioneering work of Le Cam (1973), Birge (1983, 1986), Wong and Shen (1995), Yang and Barron (1999), we determine the exact (up to constants) minimax rate over any convex density class. This work thus extends these known results by demonstrating that the local metric entropy of the density class always captures the minimax optimal rates under such settings. Our bounds provide a unifying perspective across both parametric and nonparametric convex density classes, under weaker assumptions on the richness of the density class than previously considered. Our proposed `multistage sieve' MLE applies to any such convex density class. We further demonstrate that this estimator is also adaptive to the true underlying density of interest. We apply our risk bounds to rederive known minimax rates including bounded total variation, and Holder density classes. We further illustrate the utility of the result by deriving upper bounds for less studied classes, e.g., convex mixture of densities.

math.ST

Robust Signal Detection with Quadratically Convex Orthosymmetric Constraints

This paper studies the problem of robust signal detection in Gaussian noise under quadratically convex orthosymmetric (QCO) constraints. We consider a minimax testing framework where the signal belongs to a QCO set and is separated from zero in Euclidean norm, while an adversary is allowed to arbitrarily corrupt a fraction $\epsilon $ of the samples. We establish the minimax separation radius between the null and alternative purely in terms of the constraint geometry, sample size, corruption rate, and noise scale. Our analysis argues that the Kolmogorov widths of the constraint set play a central role in determining the detection limits, paralleling to classic results in estimation problem. The derived lower bounds exhibit phase transitions with respect to the corruption rate and confirm that robust testing is statistically easier than robust estimation. While the information-theoretic upper bound is achieved by a computationally intractable test, we develop a polynomial-time algorithm that achieves the minimax lower bound up to logarithmic factors. Unlike prior work, our algorithm handles signals of arbitrary Euclidean length while respecting the QCO constraints. Finally, we extend these results to the robust $\ell _{p}$ norm testing for $1 \le p < 2$.

math.ST

Nearly Minimax Optimal Wasserstein Conditional Independence Testing

This paper is concerned with minimax conditional independence testing. In contrast to some previous works on the topic, which use the total variation distance to separate the null from the alternative, here we use the Wasserstein distance. In addition, we impose Wasserstein smoothness conditions which on bounded domains are weaker than the corresponding total variation smoothness imposed, for instance, by Neykov et al. [2021]. This added flexibility expands the distributions which are allowed under the null and the alternative to include distributions which may contain point masses for instance. We characterize the optimal rate of the critical radius of testing up to logarithmic factors. Our test statistic which nearly achieves the optimal critical radius is novel, and can be thought of as a weighted multi-resolution version of the U-statistic studied by Neykov et al. [2021].

math.ST

Non-Asymptotic Bounds for the $\ell_{\infty}$ Estimator in Linear Regression with Uniform Noise

The Chebyshev or $\ell_{\infty}$ estimator is an unconventional alternative to the ordinary least squares in solving linear regressions. It is defined as the minimizer of the $\ell_{\infty}$ objective function \begin{align*} \hat{\boldsymbolβ} := \arg\min_{\boldsymbolβ} \|\boldsymbol{Y} - \mathbf{X}\boldsymbolβ\|_{\infty}. \end{align*} The asymptotic distribution of the Chebyshev estimator under fixed number of covariates was recently studied (Knight, 2020), yet finite sample guarantees and generalizations to high-dimensional settings remain open. In this paper, we develop non-asymptotic upper bounds on the estimation error $\|\hat{\boldsymbolβ}-\boldsymbolβ^*\|_2$ for a Chebyshev estimator $\hat{\boldsymbolβ}$, in a regression setting with uniformly distributed noise $\varepsilon_i\sim U([-a,a])$ where $a$ is either known or unknown. With relatively mild assumptions on the (random) design matrix $\mathbf{X}$, we can bound the error rate by $\frac{C_p}{n}$ with high probability, for some constant $C_p$ depending on the dimension $p$ and the law of the design. Furthermore, we illustrate that there exist designs for which the Chebyshev estimator is (nearly) minimax optimal. On the other hand we also argue that there exist designs for which this estimator behaves sub-optimally in terms of the constant $C_p$'s dependence on $p$. In addition we show that "Chebyshev's LASSO" has advantages over the regular LASSO in high dimensional situations, provided that the noise is uniform. Specifically, we argue that it achieves a much faster rate of estimation under certain assumptions on the growth rate of the sparsity level and the ambient dimension with respect to the sample size.

math.ST

A New Perspective on Debiasing Linear Regressions

In this paper, we propose an abstract procedure for debiasing constrained or regularized potentially high-dimensional linear models. It is elementary to show that the proposed procedure can produce $\frac{1}{\sqrt{n}}$-confidence intervals for individual coordinates (or even bounded contrasts) in models with unknown covariance, provided that the covariance has bounded spectrum. While the proof of the statistical guarantees of our procedure is simple, its implementation requires more care due to the complexity of the optimization programs we need to solve. We spend the bulk of this paper giving examples in which the proposed algorithm can be implemented in practice. One fairly general class of instances which are amenable to applications of our procedure include convex constrained least squares. We are able to translate the procedure to an abstract algorithm over this class of models, and we give concrete examples where efficient polynomial time methods for debiasing exist. Those include the constrained version of the group LASSO, regression under monotone constraints, regression with positive monotone constraints and non-negative least squares. We also demonstrate that our method can debias Minkowski gauge selectors such as the ones proposed by Cai et al. (2016) under a certain condition. This solves an open problem posed by Cai et al. (2016) on how to debias such selectors when the covariance is unknown. In addition, we show that our abstract procedure can be applied to efficiently debias group LASSO, SLOPE and square-root SLOPE, among other popular regularized procedures under certain assumptions. We provide thorough simulation results in support of our theoretical findings.

stat.ME

On the minimax rate of the Gaussian sequence model under bounded convex constraints

We determine the exact minimax rate of a Gaussian sequence model under bounded convex constraints, purely in terms of the local geometry of the given constraint set $K$. Our main result shows that the minimax risk (up to constant factors) under the squared $\ell_2$ loss is given by $ε^{*2} \wedge \operatorname{diam}(K)^2$ with \begin{align*} ε^* = \sup \bigg\{ε: \frac{ε^2}{σ^2} \leq \log M^{\operatorname{loc}}(ε)\bigg\}, \end{align*} where $\log M^{\operatorname{loc}}(ε)$ denotes the local entropy of the set $K$, and $σ^2$ is the variance of the noise. We utilize our abstract result to re-derive known minimax rates for some special sets $K$ such as hyperrectangles, ellipses, and more generally quadratically convex orthosymmetric sets. Finally, we extend our results to the unbounded case with known $σ^2$ to show that the minimax rate in that case is $ε^{*2}$.

math.ST