Searcharxiv⌕ Search

arXiv subjects

Roberto I. Oliveira

Publications and source records attributed to Roberto I. Oliveira.

At least 19 recordsLinked to original sources

Robust dimension-free estimation of simple random tensors: optimal guarantees under heavy tails and adversarial contamination

We study robust estimation of simple random tensors of arbitrary order $q\in\mathbb{N}$ under finite-moment assumptions and adversarial contamination. We propose the first robust estimator achieving near-optimal dimension-free statistical rates in this setting. The estimator attains the near-optimal corruption rate whenever $p\ge2q$ moments are finite and continues to provide nontrivial guarantees throughout the weak-moment regime $q\le p\le2q$. Being based on directional trimmed means and minimax aggregation, our estimator is adaptive to $p$ and upper bounds on hypercontractive constants without resorting to interval-intersection procedures. Our analysis extends the trimmed-mean framework underlying recent advances in robust mean and covariance estimation to arbitrary tensor order. In particular, we establish concentration inequalities for higher-order counting and truncated empirical multi-vector product processes. We believe these inequalities could be of independent interest beyond the present application, including algorithmic robust estimation.

math.ST↗

Trimmed sample means for robust uniform mean estimation and regression

It is well-known that trimmed sample means are robust against heavy tails and data contamination. This paper analyzes the performance of trimmed means and related methods in two novel contexts. The first one consists of estimating expectations of functions in a given family, with uniform error bounds; this is closely related to the problem of estimating the mean of a random vector under a general norm. The second problem considered is that of regression with quadratic loss. In both cases, trimmed-mean-based estimators are the first to obtain optimal dependence on the (adversarial) contamination level. Moreover, they also match or improve upon the state of the art in terms of heavy tails. Experiments with synthetic data show that a natural ``trimmed mean linear regression'' method often performs better than both ordinary least squares and alternative methods based on median-of-means.

math.ST↗

Finite-sample properties of the trimmed mean

The trimmed mean of $n$ scalar random variables from a distribution $P$ is the variant of the standard sample mean where the $k$ smallest and $k$ largest values in the sample are discarded for some parameter $k$. In this paper, we look at the finite-sample properties of the trimmed mean as an estimator for the mean of $P$. Assuming finite variance, we prove that the trimmed mean is ``sub-Gaussian'' in the sense of achieving Gaussian-type concentration around the mean. Under slightly stronger assumptions, we show the left and right tails of the trimmed mean satisfy a strong ratio-type approximation by the corresponding Gaussian tail, even for very small probabilities of the order $e^{-n^c}$ for some $c>0$. In the more challenging setting of weaker moment assumptions and adversarial sample contamination, we prove that the trimmed mean is minimax-optimal up to constants.

math.ST↗

Split Conformal Prediction and Non-Exchangeable Data

Split conformal prediction (CP) is arguably the most popular CP method for uncertainty quantification, enjoying both academic interest and widespread deployment. However, the original theoretical analysis of split CP makes the crucial assumption of data exchangeability, which hinders many real-world applications. In this paper, we present a novel theoretical framework based on concentration inequalities and decoupling properties of the data, proving that split CP remains valid for many non-exchangeable processes by adding a small coverage penalty. Through experiments with both real and synthetic data, we show that our theoretical results translate to good empirical performance under non-exchangeability, e.g., for time series and spatiotemporal data. Compared to recent conformal algorithms designed to counter specific exchangeability violations, we show that split CP is competitive in terms of coverage and interval size, with the benefit of being extremely simple and orders of magnitude faster than alternatives.

math.ST↗

Improved covariance estimation: optimal robustness and sub-Gaussian guarantees under heavy tails

We present an estimator of the covariance matrix $Σ$ of random $d$-dimensional vector from an i.i.d. sample of size $n$. Our sole assumption is that this vector satisfies a bounded $L^p-L^2$ moment assumption over its one-dimensional marginals, for some $p\geq 4$. Given this, we show that $Σ$ can be estimated from the sample with the same high-probability error rates that the sample covariance matrix achieves in the case of Gaussian data. This holds even though we allow for very general distributions that may not have moments of order $>p$. Moreover, our estimator can be made to be optimally robust to adversarial contamination. This result improves the recent contributions by Mendelson and Zhivotovskiy and Catoni and Giulini, and matches parallel work by Abdalla and Zhivotovskiy (the exact relationship with this last work is described in the paper).

math.ST↗

Deep Hashing via Householder Quantization

Hashing is at the heart of large-scale image similarity search, and recent methods have been substantially improved through deep learning techniques. Such algorithms typically learn continuous embeddings of the data. To avoid a subsequent costly binarization step, a common solution is to employ loss functions that combine a similarity learning term (to ensure similar images are grouped to nearby embeddings) and a quantization penalty term (to ensure that the embedding entries are close to binarized entries, e.g., -1 or 1). Still, the interaction between these two terms can make learning harder and the embeddings worse. We propose an alternative quantization strategy that decomposes the learning problem in two stages: first, perform similarity learning over the embedding space with no quantization; second, find an optimal orthogonal transformation of the embeddings so each coordinate of the embedding is close to its sign, and then quantize the transformed embedding through the sign function. In the second step, we parametrize orthogonal transformations using Householder matrices to efficiently leverage stochastic gradient descent. Since similarity measures are usually invariant under orthogonal transformations, this quantization strategy comes at no cost in terms of performance. The resulting algorithm is unsupervised, fast, hyperparameter-free and can be run on top of any existing deep hashing or metric learning algorithm. We provide extensive experimental results showing that this approach leads to state-of-the-art performance on widely used image datasets, and, unlike other quantization strategies, brings consistent improvements in performance to existing deep hashing algorithms.

cs.CV↗

Sample average approximation with heavier tails II: localization in stochastic convex optimization and persistence results for the Lasso

``Localization'' has proven to be a valuable tool in the Statistical Learning literature as it allows sharp risk bounds in terms of the problem geometry. Localized bounds seem to be much less exploited in the Stochastic Optimization literature. In addition, there is an obvious interest in both communities in obtaining risk bounds that require weak moment assumptions or ``heavier-tails''. In this work we use a localization toolbox to derive risk bounds in two specific applications. The first is in portfolio risk minimization with conditional value-at-risk constraints. We consider a setting where, among all assets with high returns, there is a portion of dimension $g$, unknown to the investor, that has significant less risk than the other remaining portion. Our rates for the SAA problem show that ``risk inflation'', caused by a multiplicative factor, affects the statistical rate only via a term proportional to $g$. As the ``normalized risk'' increases, the contribution in the rate from the extrinsic dimension diminishes while the dependence on $g$ is kept fixed. Localization is a key tool to show this property. As a second application of our localization toolbox, we obtain sharp oracle inequalities for least-squares estimators with a Lasso-type constraint under weak moment assumptions. One main consequence of these inequalities is to obtain \emph{persistence}, as posed by Greenshtein and Ritov, with covariates having heavier tails. This gives improvements in prior work of Bartlett, Mendelson and Neeman.

math.OC↗

A spectral least-squares-type method for heavy-tailed corrupted regression with unknown covariance \& heterogeneous noise

We revisit heavy-tailed corrupted least-squares linear regression assuming to have a corrupted $n$-sized label-feature sample of at most $εn$ arbitrary outliers. We wish to estimate a $p$-dimensional parameter $b^*$ given such sample of a label-feature pair $(y,x)$ satisfying $y=\langle x,b^*\rangle+ξ$ with heavy-tailed $(x,ξ)$. We only assume $x$ is $L^4-L^2$ hypercontractive with constant $L>0$ and has covariance matrix $Σ$ with minimum eigenvalue $1/μ^2>0$ and bounded condition number $κ>0$. The noise $ξ$ can be arbitrarily dependent on $x$ and nonsymmetric as long as $ξx$ has finite covariance matrix $Ξ$. We propose a near-optimal computationally tractable estimator, based on the power method, assuming no knowledge on $(Σ,Ξ)$ nor the operator norm of $Ξ$. With probability at least $1-δ$, our proposed estimator attains the statistical rate $μ^2\VertΞ\Vert^{1/2}(\frac{p}{n}+\frac{\log(1/δ)}{n}+ε)^{1/2}$ and breakdown-point $ε\lesssim\frac{1}{L^4κ^2}$, both optimal in the $\ell_2$-norm, assuming the near-optimal minimum sample size $L^4κ^2(p\log p + \log(1/δ))\lesssim n$, up to a log factor. To the best of our knowledge, this is the first computationally tractable algorithm satisfying simultaneously all the mentioned properties. Our estimator is based on a two-stage Multiplicative Weight Update algorithm. The first stage estimates a descent direction $\hat v$ with respect to the (unknown) pre-conditioned inner product $\langleΣ(\cdot),\cdot\rangle$. The second stage estimate the descent direction $Σ\hat v$ with respect to the (known) inner product $\langle\cdot,\cdot\rangle$, without knowing nor estimating $Σ$.

math.ST↗

A proof of Sanov's Theorem via discretizations

We present an alternative proof of Sanov's theorem for Polish spaces in the weak topology that follows via discretization arguments. We combine the simpler version of Sanov's Theorem for discrete finite spaces and well chosen finite discretizations of the Polish space. The main tool in our proof is an explicit control on the rate of convergence for the approximated measures.

math.PR↗

Sample average approximation with heavier tails I: non-asymptotic bounds with weak assumptions and stochastic constraints

We derive new and improved non-asymptotic deviation inequalities for the sample average approximation (SAA) of an optimization problem. Our results give strong error probability bounds that are "sub-Gaussian"~even when the randomness of the problem is fairly heavy tailed. Additionally, we obtain good (often optimal) dependence on the sample size and geometrical parameters of the problem. Finally, we allow for random constraints on the SAA and unbounded feasible sets, which also do not seem to have been considered before in the non-asymptotic literature. Our proofs combine different ideas of potential independent interest: an adaptation of Talagrand's "generic chaining"~bound for sub-Gaussian processes; "localization"~ideas from the Statistical Learning literature; and the use of standard conditions in Optimization (metric regularity, Slater-type conditions) to control fluctuations of the feasible set.

math.OC↗

Interacting diffusions on sparse graphs: hydrodynamics from local weak limits

We prove limit theorems for systems of interacting diffusions on sparse graphs. For example, we deduce a hydrodynamic limit and the propagation of chaos property for the stochastic Kuramoto model with interactions determined by Erdős-Rényi graphs with constant mean degree. The limiting object is related to a potentially infinite system of SDEs defined over a Galton-Watson tree. Our theorems apply more generally, when the sequence of graphs ("decorated" with edge and vertex parameters) converges in the local weak sense. Our main technical result is a locality estimate bounding the influence of far-away diffusions on one another. We also numerically explore the emergence of synchronization phenomena on Galton-Watson random trees, observing rich phase transitions from synchronized to desynchronized activity among nodes at different distances from the root.

math.PR↗

A mean-field limit for certain deep neural networks

Understanding deep neural networks (DNNs) is a key challenge in the theory of machine learning, with potential applications to the many fields where DNNs have been successfully used. This article presents a scaling limit for a DNN being trained by stochastic gradient descent. Our networks have a fixed (but arbitrary) number $L\geq 2$ of inner layers; $N\gg 1$ neurons per layer; full connections between layers; and fixed weights (or "random features" that are not trained) near the input and output. Our results describe the evolution of the DNN during training in the limit when $N\to +\infty$, which we relate to a mean field model of McKean-Vlasov type. Specifically, we show that network weights are approximated by certain "ideal particles" whose distribution and dependencies are described by the mean-field model. A key part of the proof is to show existence and uniqueness for our McKean-Vlasov problem, which does not seem to be amenable to existing theory. Our paper extends previous work on the $L=1$ case by Mei, Montanari and Nguyen; Rotskoff and Vanden-Eijnden; and Sirignano and Spiliopoulos. We also complement recent independent work on $L>1$ by Sirignano and Spiliopoulos (who consider a less natural scaling limit) and Nguyen (who nonrigorously derives similar results).

math.ST↗

Building your path to escape from home

Random walks on dynamic graphs have received increasingly more attention from different academic communities over the last decade. Despite the relatively large literature, little is known about random walks that construct the graph where they walk while moving around. In this paper we study one of the simplest conceivable discrete time models of this kind, which works as follows: before every walker step, with probability $p$ a new leaf is added to the vertex currently occupied by the walker. The model grows trees and we call it the Bernoulli Growth Random Walk (BGRW). We show that the BGRW walker is transient and has a well-defined linear speed $c(p)>0$ for any $0<p\leq 1$. We also show that the tree as seen by the walker converges (in a suitable sense) to a random tree that is one-ended. Some natural open problems about this tree and variants of our model are collected at the end of the paper.

math.PR↗

Interacting diffusions on random graphs with diverging degrees: hydrodynamics and large deviations

We consider systems of mean-field interacting diffusions, where the pairwise interaction structure is described by a sparse (and potentially inhomogeneous) random graph. Examples include the stochastic Kuramoto model with pairwise interactions given by an Erdős-Rényi graph. Our problem is to compare the bulk behavior of such systems with that of corresponding systems with dense nonrandom interactions. For a broad class of interaction functions, we find the optimal sparsity condition that implies that the two systems have the same hydrodynamic limit, which is given by a McKean-Vlasov diffusion. Moreover, we also prove matching behavior of the two systems at the level of large deviations. Our results extend classical results of dai Pra and den Hollander and provide the first examples of LDPs for systems with sparse random interactions.

math.PR↗

A high dimensional Central Limit Theorem for martingales, with applications to context tree models

We establish a central limit theorem for (a sequence of) multivariate martingales which dimension potentially grows with the length $n$ of the martingale. A consequence of the results are Gaussian couplings and a multiplier bootstrap for the maximum of a multivariate martingale whose dimensionality $d$ can be as large as $e^{n^c}$ for some $c>0$. We also develop new anti-concentration bounds for the maximum component of a high-dimensional Gaussian vector, which we believe is of independent interest. The results are applicable to a variety of settings. We fully develop its use to the estimation of context tree models (or variable length Markov chains) for discrete stationary time series. Specifically, we provide a bootstrap-based rule to tune several regularization parameters in a theoretically valid Lepski-type method. Such bootstrap-based approach accounts for the correlation structure and leads to potentially smaller penalty choices, which in turn improve the estimation of the transition probabilities.

math.ST↗

Estimating graph parameters with random walks

An algorithm observes the trajectories of random walks over an unknown graph $G$, starting from the same vertex $x$, as well as the degrees along the trajectories. For all finite connected graphs, one can estimate the number of edges $m$ up to a bounded factor in $O\left(t_{\mathrm{rel}}^{3/4}\sqrt{m/d}\right)$ steps, where $t_{\mathrm{rel}}$ is the relaxation time of the lazy random walk on $G$ and $d$ is the minimum degree in $G$. Alternatively, $m$ can be estimated in $O\left(t_{\mathrm{unif}} +t_{\mathrm{rel}}^{5/6}\sqrt{n}\right)$, where $n$ is the number of vertices and $t_{\mathrm{unif}}$ is the uniform mixing time on $G$. The number of vertices $n$ can then be estimated up to a bounded factor in an additional $O\left(t_{\mathrm{unif}}\frac{m}{n}\right)$ steps. Our algorithms are based on counting the number of intersections of random walk paths $X,Y$, i.e. the number of pairs $(t,s)$ such that $X_t=Y_s$. This improves on previous estimates which only consider collisions (i.e., times $t$ with $X_t=Y_t$). We also show that the complexity of our algorithms is optimal, even when restricting to graphs with a prescribed relaxation time. Finally, we show that, given either $m$ or the mixing time of $G$, we can compute the "other parameter" with a self-stopping algorithm.

math.ST↗

Random walks on graphs: new bounds on hitting, meeting, coalescing and returning

We prove new results on lazy random walks on finite graphs. To start, we obtain new estimates on return probabilities $P^t(x,x)$ and the maximum expected hitting time $t_{\rm hit}$, both in terms of the relaxation time. We also prove a discrete-time version of the first-named author's ``Meeting time lemma"~ that bounds the probability of random walk hitting a deterministic trajectory in terms of hitting times of static vertices. The meeting time result is then used to bound the expected full coalescence time of multiple random walks over a graph. This last theorem is a discrete-time version of a result by the first-named author, which had been previously conjectured by Aldous and Fill. Our bounds improve on recent results by Lyons and Oveis-Gharan; Kanade et al; and (in certain regimes) Cooper et al.

math.PR↗

Concentration in the Generalized Chinese Restaurant Process

The Generalized Chinese Restaurant Process (GCRP) describes a sequence of exchangeable random partitions of the numbers $\{1,\dots,n\}$. This process is related to the Ewens sampling model in Genetics and to Bayesian nonparametric methods such as topic models. In this paper, we study the GCRP in a regime where the number of parts grows like $n^α$ with $α>0$. We prove a non-asymptotic concentration result for the number of parts of size $k=o(n^{α/(2α+4)}/(\log n)^{1/(2+α)})$. In particular, we show that these random variables concentrate around $c_{k}\,V_*\,n^α$ where $V_*\,n^α$ is the asymptotic number of parts and $c_k\approx k^{-(1+α)}$ is a positive value depending on $k$. We also obtain finite-$n$ bounds for the total number of parts. Our theorems complement asymptotic statements by Pitman and more recent results on large and moderate deviations by Favaro, Feng and Gao.

math.PR↗