SearcharxivSearch

arXiv subjects

Fedor Noskov

Publications and source records attributed to Fedor Noskov.

15 recordsLinked to original sources

Density Estimation on Compact Manifolds under Intrinsic Spectral Block Variation

We introduce an intrinsic spectral sparsity model for nonparametric density estimation on compact connected Riemannian manifolds. Instead of penalizing coefficients in an arbitrarily chosen Laplace--Beltrami eigenbasis, we group each complete eigenspace and measure the Hilbert norm of its spectral component. The resulting block-variation space is basis independent and isometry invariant. We establish its structural, atomic, and nonlinear approximation properties and clarify its relation to Sobolev, Besov, and coefficientwise spectral $\ell^1$ classes. We then construct a coordinate-free block-shrinkage estimator and prove a nonasymptotic signal-dependent $L^2$-oracle inequality that adapts to the unknown set of detectable eigenspaces. Under polynomial spectral growth, the risk theory separates the number of spectral blocks from their multiplicities and exhibits two regimes: one driven by a single high-dimensional eigenspace and the other by cumulative spectral complexity. Under matching spectral-growth and nondegeneracy assumptions, corresponding minimax lower bounds show that this multiplicity dependence is intrinsic, with sharp consequences for spheres and the rotation group $SO(3)$. Finally, we develop a positive, normalized, block-penalized exponential spectral sieve for log-densities and derive likelihood oracle inequalities together with expected Kullback--Leibler, Hellinger, and $L^2$ risk bounds. The resulting framework provides a geometry-respecting theory of sparse density estimation that remains invariant under changes of eigenbasis.

math.ST

Forbidding just one intersection for short integer sequences

In this paper, we study the famous Erd\H{o}s--S\'os forbidden intersection problem for words over an alphabet of size $m$: what is the maximal size of a subfamily $\mathcal{F}$ of $[m]^n$ that does not contain two vectors $x, y$ coinciding on exactly $t - 1$ coordinates? We answer this question provided $m \ge \operatorname{poly}(t)$ and $n \ge \operatorname{poly}(t)$ for some polynomial function $\operatorname{poly}(\cdot)$ of $t$, greatly extending the recent result of Keevash, Lifshitz, Long and Minzer. Our proof combines some of the recently developed methods in extremal combinatorics, including the spread approximation technique of Kupavskii and Zakharov and the hypercontractivity approach developed in a series of works by Keevash, Keller, Lifshitz, Long, Marcus and Minzer.

math.CO

Exact results and the structure of extremal families for the Duke--Erd\H{o}s forbidden sunflower problem

In 1977, Duke and Erd\H{o}s asked the following general question: What is the largest size of a family $\mathtt{F} \subset \binom{[n]}{k}$ that does not contain a sunflower with $s$ petals and core of size exactly $t - 1$? This problem is closely related to the famous Erd\H{o}s--Rado sunflower problem of determining the size $\phi(s,t)$ of the largest $t$-uniform family with no $s$-sunflower. In this paper, we answer this question exactly for $t=2$, odd $s$ and $k\ge 5$, provided $n$ is large enough. Previously, the only know exact extremal result on this problem was due to Chung and Frankl from 1987. One of the important ingredients for the proof that we obtained is a stability result for the Duke--Erd\H{o}s problem, which was previously not known, mostly due to our lack of understanding of the behaviour of $\phi(s,t)$. For large $k$ and $n$ we in fact manage to reduce the Duke--Erd\H{o}s problem to an Erd\H{o}s--Rado-like problem which depends on $t$ and $s$ only. In particular, we get a good understanding of the structure of extremal families for the Duke--Erd\H{o}s problem in terms of the Erd\H{o}s--Rado problem. Previously, a much looser variant of this connection (only in terms of the sizes, rather than the structure, of respective extremal families) was established in a seminal work of Frankl and F\"uredi from 1987.

math.CO

Dimension-free Bounds for Covariance Estimation with Tensor-Train Structure

We consider a problem of covariance estimation from a sample of i.i.d. high-dimensional random vectors. To avoid the curse of dimensionality, we impose an additional assumption on the structure of the covariance matrix $\Sigma$. To be more precise, we study the case when $\Sigma$ can be approximated by a sum of double Kronecker products of smaller matrices in a tensor train (TT) format. Our setup naturally extends widely known Kronecker sum and CANDECOMP/PARAFAC models but admits richer interaction across modes. We suggest an iterative polynomial time algorithm based on TT-SVD and higher-order orthogonal iteration (HOOI) adapted to Tucker-2 hybrid structure. We derive non-asymptotic dimension-free bounds on the accuracy of covariance estimation taking into account hidden Kronecker product and tensor train structures. The efficiency of our approach is illustrated with numerical experiments.

math.ST

Low-Rank Graphon Estimation: Theory and Applications to Graphon Games

We study low-rank estimation of an unknown sparse graphon from sampled network data under operator-norm loss, motivated by targeted interventions in graphon games. Starting from the observed adjacency matrix, we construct low-rank surrogates by singular value thresholding and, for smooth graphons, by block averaging followed by thresholding. We obtain non-asymptotic bounds on both the operator-norm error and the rank of the resulting estimator for stochastic block model, H\"older, and analytic graphons, and we complement these results with minimax lower bounds showing that the rates are essentially sharp for these classes. Our analysis highlights that low rank is valuable here primarily for computation: while it does not improve the minimax operator-norm rate, it yields operator-norm accurate surrogates with substantially smaller rank. We then apply these estimators to linear-quadratic graphon games and derive non-asymptotic stability bounds showing that the welfare loss incurred by using an estimated graphon is controlled by the operator-norm perturbation. This yields near-optimal guarantees for targeted interventions computed from the estimated graphon, together with substantial computational savings. For zero baseline heterogeneity and under a spectral-gap condition, we also establish matching lower bounds for intervention regret. Numerical experiments illustrate the trade-off between statistical accuracy, retained rank, and runtime.

math.ST

Dimension-free bounds in high-dimensional linear regression via error-in-operator approach

We consider a problem of high-dimensional linear regression with random design. We suggest a novel approach referred to as error-in-operator which does not estimate the design covariance $\Sigma$ directly but incorporates it into empirical risk minimization. We provide an expansion of the excess prediction risk and derive non-asymptotic dimension-free bounds on the leading term and the remainder. This helps us to show that auxiliary variables do not increase the effective dimension of the problem, provided that parameters of the procedure are tuned properly. We also discuss computational aspects of our method and illustrate its performance with numerical experiments.

math.ST

Linear dependencies, polynomial factors in the Duke--Erd\H os forbidden sunflower problem

We call a family of $s$ sets $\{F_1, \ldots, F_s\}$ a \textit{sunflower with $s$ petals} if, for any distinct $i, j \in [s]$, one has $F_i \cap F_j = \cap_{u = 1}^s F_u$. The set $C = \cap_{u = 1}^s F_u$ is called the {\it core} of the sunflower. It is a classical result of Erd\H os and Rado that there is a function $\phi(s,k)$ such that any family of $k$-element sets contains a sunflower with $s$ petals. In 1977, Duke and Erd\H os asked for the size of the largest family $\mathcal{F}\subset{[n]\choose k}$ that contains no sunflower with $s$ petals and core of size $t-1$. In 1987, Frankl and F\" uredi asymptotically solved this problem for $k\ge 2t+1$ and $n>n_0(s,k)$. This paper is one of the pinnacles of the so-called Delta-system method. In this paper, we extend the result of Frankl and F\"uredi to a much broader range of parameters: $n>f_0(s,t) k$ with $f_0(s,t)$ polynomial in $s$ and $t$. We also extend this result to other domains, such as $[n]^k$ and ${n\choose k/w}^w$ and obtain even stronger and more general results for forbidden sunflowers with core at most $t-1$ (including results for families of permutations and subfamilies of the $k$-th layer in a simplicial complex). The methods of the paper, among other things, combine the spread approximation technique, introduced by Zakharov and the first author, with the Delta-system approach of Frankl and F\"uredi and the hypercontractivity approach for global functions, developed by Keller, Lifshitz and coauthors. Previous works in extremal set theory relied on at most one of these methods. Creating such a unified approach was one of the goals for the paper.

math.CO

Efficient Conformal Prediction under Data Heterogeneity

Conformal Prediction (CP) stands out as a robust framework for uncertainty quantification, which is crucial for ensuring the reliability of predictions. However, common CP methods heavily rely on data exchangeability, a condition often violated in practice. Existing approaches for tackling non-exchangeability lead to methods that are not computable beyond the simplest examples. This work introduces a new efficient approach to CP that produces provably valid confidence sets for fairly general non-exchangeable data distributions. We illustrate the general theory with applications to the challenging setting of federated learning under data heterogeneity between agents. Our method allows constructing provably valid personalized prediction sets for agents in a fully federated way. The effectiveness of the proposed method is demonstrated in a series of experiments on real-world datasets.

stat.ML

Selective Nonparametric Regression via Testing

Prediction with the possibility of abstention (or selective prediction) is an important problem for error-critical machine learning applications. While well-studied in the classification setup, selective approaches to regression are much less developed. In this work, we consider the nonparametric heteroskedastic regression problem and develop an abstention procedure via testing the hypothesis on the value of the conditional variance at a given point. Unlike existing methods, the proposed one allows to account not only for the value of the variance itself but also for the uncertainty of the corresponding variance predictor. We prove non-asymptotic bounds on the risk of the resulting estimator and show the existence of several different convergence regimes. Theoretical analysis is illustrated with a series of experiments on simulated and real-world data.

stat.ML

Octopuses in the Boolean cube: families with pairwise small intersections, part II

The problem we consider originally arises from 2-level polytope theory. This class of polytopes generalizes a number of other polytope families. One of the important questions in this filed can be formulated as follows: is it true for a $d$-dimensional 2-level polytope that the product of the number of its vertices and the number of its $d-1$ dimensional facets is bounded by $d2^{d - 1}$? Recently, Kupavskii and Weltge~\cite{Kupavskii2020} settled this question in positive. A key element in their proof is a more general result for families of vectors in $\mathbb{R}^d$ such that the scalar product between any two vectors from different families is either $0$ or $1$. Peter Frankl noted that, when restricted to the Boolean cube, the solution boils down to an elegant application of the Harris--Kleitman correlation inequality. Meanwhile, this problem becomes much more sophisticated when we consider several families. Let $\mathcal{F}_1, \ldots, \mathcal{F}_\ell$ be families of subsets of $\{1, \ldots, n\}$. We suppose that for distinct $k, k'$ and arbitrary $F_1 \in \mathcal{F}_{k}, F_2 \in \mathcal{F}_{k'}$ we have $|F_1 \cap F_2|\leqslant m.$ We are interested in the maximal value of $|\mathcal{F}_1|\ldots |\mathcal{F}_\ell|$ and the structure of the extremal example. In the previous paper on the topic, the authors found the asymptotics of this product for constant $\ell$ and $m$ as $n$ tends to infinity. However, the possible structure of the families from the extremal example turned out to be very complicated. In this paper, we obtain a strong structural result for the extremal families.

math.CO

Sharper dimension-free bounds on the Frobenius distance between sample covariance and its expectation

We study properties of a sample covariance estimate $\widehat \Sigma$ given a finite sample of $n$ i.i.d. centered random elements in $\R^d$ with the covariance matrix $\Sigma$. We derive dimension-free bounds on the squared Frobenius norm of $(\widehat\Sigma - \Sigma)$ under reasonable assumptions. For instance, we show that $\smash{\|\widehat\Sigma - \Sigma\|_{\rm F}^2}$ differs from its expectation by at most $\smash{\mathcal O({\rm{Tr}}(\Sigma^2) / n)}$ with overwhelming probability, which is a significant improvement over the existing results. This allows us to establish the concentration phenomenon for the squared Frobenius distance between the covariance and its empirical counterpart in the case of moderately large effective rank of $\Sigma$.

math.PR

Optimal Noise Reduction in Dense Mixed-Membership Stochastic Block Models under Diverging Spiked Eigenvalues Condition

Community detection is one of the most critical problems in modern network science. Its applications can be found in various fields, from protein modeling to social network analysis. Recently, many papers appeared studying the problem of overlapping community detection, where each node of a network may belong to several communities. In this work, we consider Mixed-Membership Stochastic Block Model (MMSB) first proposed by Airoldi et al. MMSB provides quite a general setting for modeling overlapping community structure in graphs. The central question of this paper is to reconstruct relations between communities given an observed network. We compare different approaches and establish the minimax lower bound on the estimation error. Then, we propose a new estimator that matches this lower bound. Theoretical results are proved under fairly general conditions on the considered model. Finally, we illustrate the theory in a series of experiments.

stat.ML

Nonparametric Uncertainty Quantification for Single Deterministic Neural Network

This paper proposes a fast and scalable method for uncertainty quantification of machine learning models' predictions. First, we show the principled way to measure the uncertainty of predictions for a classifier based on Nadaraya-Watson's nonparametric estimate of the conditional label distribution. Importantly, the proposed approach allows to disentangle explicitly aleatoric and epistemic uncertainties. The resulting method works directly in the feature space. However, one can apply it to any neural network by considering an embedding of the data induced by the network. We demonstrate the strong performance of the method in uncertainty estimation tasks on text classification problems and a variety of real-world image datasets, such as MNIST, SVHN, CIFAR-100 and several versions of ImageNet.

stat.ML

Octopuses in the Boolean cube: families with pairwise small intersections, part I

Let $\mathcal F_1, \ldots, \mathcal F_\ell$ be families of subsets of $\{1, \ldots, n\}$. Suppose that for distinct $k, k'$ and arbitrary $F_1 \in \mathcal F_{k}, F_2 \in \mathcal F_{k'}$ we have $|F_1 \cap F_2|\le m.$ What is the maximal value of $|\mathcal F_1|\ldots |\mathcal F_\ell|$? In this work we find the asymptotic of this product as $n$ tends to infinity for constant $\ell$ and~$m$. This question is related to a conjecture of Bohn et al. that arose in the 2-level polytope theory and asked for the largest product of the number of facets and vertices in a two-level polytope. This conjecture was recently resolved by Weltge and the first author. The main result can be rephrased in terms of colorings. We give an asymptotic answer to the following question. Given an edge coloring of a complete $m$-uniform hypergraph into $\ell$ colors, what is the maximum of $\prod M_i$, where $M_i$ is the number of monochromatic cliques in $i$-th color?

math.CO

Around power law for PageRank components in Buckley-Osthus model of web graph

In the paper we investigate power law for PageRank components for the Buckley-Osthus model for web graph. We compare different numerical methods for PageRank calculation. With the best method we do a lot of numerical experiments. These experiments confirm the hypothesis about power law. At the end we discuss real model of web-ranking based on the classical PageRank approach.

math.OC