SearcharxivSearch

arXiv subjects

Leonardo Moreno

Publications and source records attributed to Leonardo Moreno.

At least 19 recordsLinked to original sources

Conformal Robust Set Estimation

Conformal prediction provides finite-sample, distribution-free coverage under exchangeability, but standard constructions may lack robustness in the presence of outliers or heavy tails. We propose a robust conformal method based on a non-conformity score defined as the half-mass radius around a point, equivalently the distance to its $(\lfloor n/2\rfloor+1)$-nearest neighbour. We show that the resulting conformal regions are marginally valid for any sample size and converge in probability to a robust population central set defined through a distance-to-a-measure functional. Under mild regularity conditions, we establish exponential concentration and tail bounds that quantify the deviation between the empirical conformal region and its population counterpart. These results provide a probabilistic justification for using robust geometric scores in conformal prediction, even for heavy-tailed or multi-modal distributions.

math.ST

Depth based trimmed means

Robust estimation of location is a fundamental problem in statistics, particularly in scenarios where data contamination by outliers or model misspecification is a concern. In univariate settings, methods such as the sample median and trimmed means balance robustness and efficiency by mitigating the influence of extreme observations. This paper extends these robust techniques to the multivariate context through the use of data depth functions, which provide a natural means to order and rank multidimensional data. We review several depth measures and discuss their role in generalizing trimmed mean estimators beyond one dimension. Our main contributions are twofold: first, we prove the almost sure consistency of the multivariate trimmed mean estimator under mixing conditions; second, we establish a general limit distribution theorem for a broad family of depth-based estimators, encompassing popular examples such as Tukey's and projection depth. These theoretical advancements not only enhance the understanding of robust location estimation in high-dimensional settings but also offer practical guidelines for applications in areas such as machine learning, economic analysis, and financial risk assessment. A small example with simulated data is performed, varying the depth measure used and the percentage of trimmed data.

math.ST

Two statistical problems for multivariate mixture distributions

We address two important statistical problems: that of estimating mixtures of multivariate normal distributions and mixtures of $t$-distributions based on univariate projections, and that of quantifying a discrepancy between mixture distributions induced by two model-based clusterings. In the second problem, rather than introducing a direct metric on partitions, we propose a model-based distributional discrepancy between the fitted mixture distributions associated with two clusterings. The results are based on an earlier work of the authors, where it was shown that mixtures of multivariate Gaussian or $t$-distributions can be distinguished by projecting them onto a certain predetermined finite set of lines, the number of lines depending only on the total number of distributions involved and on the ambient dimension. We also compare our proposal with robust versions of the expectation-maximization method EM. In each case, we present algorithms for effecting the task, and compare them with existing methods by carrying out some simulations.

math.ST

A Cram\'er-Wold theorem for mixtures

We show how a Cram\'er-Wold theorem for a family of multivariate probability distributions can be used to generate a similar theorem for mixtures (convex combinations) of distributions drawn from the same family. Using this abstract result, we establish a Cram\'er-Wold theorem for mixtures of multivariate Gaussian distributions. According to this theorem, two such mixtures can be distinguished by projecting them onto a certain predetermined finite set of lines, the number of lines depending only on the total number Gaussian distributions involved and on the ambient dimension. A similar result is also obtained for mixtures of multivariate $t$-distributions.

math.PR

GROS: A General Robust Aggregation Strategy

A new, very general, robust procedure for combining estimators in metric spaces is introduced GROS. The method is reminiscent of the well-known median of means, as described in \cite{devroye2016sub}. Initially, the sample is divided into $K$ groups. Subsequently, an estimator is computed for each group. Finally, these $K$ estimators are combined using a robust procedure. We prove that this estimator is sub-Gaussian and we get its break-down point, in the sense of Donoho. The robust procedure involves a minimization problem on a general metric space, but we show that the same (up to a constant) sub-Gaussianity is obtained if the minimization is taken over the sample, making GROS feasible in practice. The performance of GROS is evaluated through five simulation studies: the first one focuses on classification using $k$-means, the second one on the multi-armed bandit problem, the third one on the regression problem. The fourth one is the set estimation problem under a noisy model. Lastly, we apply GROS to get a robust persistent diagram.

math.ST

On the notion of polynomial reach: a statistical application

The volume function V(t) of a compact set S\in R^d is just the Lebesgue measure of the set of points within a distance to S not larger than t. According to some classical results in geometric measure theory, the volume function turns out to be a polynomial, at least in a finite interval, under a quite intuitive, easy to interpret, sufficient condition (called ``positive reach'') which can be seen as an extension of the notion of convexity. However, many other simple sets, not fulfilling the positive reach condition, have also a polynomial volume function. To our knowledge, there is no general, simple geometric description of such sets. Still, the polynomial character of $V(t)$ has some relevant consequences since the polynomial coefficients carry some useful geometric information. In particular, the constant term is the volume of S and the first order coefficient is the boundary measure (in Minkowski's sense). This paper is focused on sets whose volume function is polynomial on some interval starting at zero, whose length (that we call ``polynomial reach'') might be unknown. Our main goal is to approximate such polynomial reach by statistical means, using only a large enough random sample of points inside S. The practical motivation is simple: when the value of the polynomial reach , or rather a lower bound for it, is approximately known, the polynomial coefficients can be estimated from the sample points by using standard methods in polynomial approximation. As a result, we get a quite general method to estimate the volume and boundary measure of the set, relying only on an inner sample of points and not requiring the use any smoothing parameter. This paper explores the theoretical and practical aspects of this idea.

math.ST

SpICE: An interpretable method for spatial data

Statistical learning methods are widely utilized in tackling complex problems due to their flexibility, good predictive performance and its ability to capture complex relationships among variables. Additionally, recently developed automatic workflows have provided a standardized approach to implementing statistical learning methods across various applications. However these tools highlight a main drawbacks of statistical learning: its lack of interpretation in their results. In the past few years an important amount of research has been focused on methods for interpreting black box models. Having interpretable statistical learning methods is relevant to have a deeper understanding of the model. In problems were spatial information is relevant, combined interpretable methods with spatial data can help to get better understanding of the problem and interpretation of the results. This paper is focused in the individual conditional expectation (ICE-plot), a model agnostic methods for interpreting statistical learning models and combined them with spatial information. ICE-plot extension is proposed where spatial information is used as restriction to define Spatial ICE curves (SpICE). Spatial ICE curves are estimated using real data in the context of an economic problem concerning property valuation in Montevideo, Uruguay. Understanding the key factors that influence property valuation is essential for decision-making, and spatial data plays a relevant role in this regard.

stat.ME

Conformal inference for regression on Riemannian Manifolds

Regression on manifolds, and, more broadly, statistics on manifolds, has garnered significant importance in recent years due to the vast number of applications for non Euclidean data. Circular data is a classic example, but so is data in the space of covariance matrices, data on the Grassmannian manifold obtained as a result of principal component analysis, among many others. In this work we investigate prediction sets for regression scenarios when the response variable, denoted by $Y$, resides in a manifold, and the covariable, denoted by $X$, lies in an Euclidean space. This extends the concepts delineated in \cite{waser14} to this novel context. Aligning with traditional principles in conformal inference, these prediction sets are distribution-free, indicating that no specific assumptions are imposed on the joint distribution of $(X,Y)$, and they maintain a non-parametric character. We prove the asymptotic almost sure convergence of the empirical version of these regions on the manifold to their population counterparts. The efficiency of this method is shown through a comprehensive simulation study and an analysis involving real-world data.

stat.ML

On standardness and the non-estimability of certain functionals of a set

Standardness is a popular assumption in the literature on set estimation. It also appears in statistical approaches to topological data analysis, where it is common to assume that the data were sampled from a probability measure that satisfies the standard assumption. Relevant results in this field, such as rates of convergence and confidence sets, depend on the standardness parameter, which in practice may be unknown. In this paper, we review the notion of standardness and its connection to other geometrical restrictions. We prove the almost sure consistency of a plug-in type estimator for the so-called standardness constant, already studied in the literature. We propose a method to correct the bias of the plug-in estimator and corroborate our theoretical findings through a small simulation study. We also show that it is not possible to determine, based on a finite sample, whether a probability measure satisfies the standard assumption.

math.ST

A quantitative Heppes Theorem and multivariate Bernoulli distributions

Using some extensions of a theorem of Heppes on finitely supported discrete probability measures, we address the problems of classification and testing based on projections. In particular, when the support of the distributions is known in advance (as for instance for multivariate Bernoulli distributions), a single suitably chosen projection determines the distribution. Several applications of these results are considered.

math.PR

A Cramér-Wold theorem for elliptical distributions

According to a well-known theorem of Cramér and Wold, if $P$ and $Q$ are two Borel probability measures on $\mathbb{R}^d$ whose projections $P_L,Q_L$ onto each line $L$ in $\mathbb{R}^d$ satisfy $P_L=Q_L$, then $P=Q$. Our main result is that, if $P$ and $Q$ are both elliptical distributions, then, to show that $P=Q$, it suffices merely to check that $P_L=Q_L$ for a certain set of $(d^2+d)/2$ lines $L$. Moreover $(d^2+d)/2$ is optimal. The class of elliptical distributions contains the Gaussian distributions as well as many other multivariate distributions of interest. Our theorem contrasts with other variants of the Cramér-Wold theorem, in that no assumption is made about the finiteness of moments of $P$ and $Q$. We use our results to derive a statistical test for equality of elliptical distributions, and carry out a small simulation study of the test, comparing it with other tests from the literature. We also give an application to learning (binary classification), again illustrated with a small simulation

math.PR

Universally consistent estimation of the reach

The reach of a set $M \subset \mathbb R^d$, also known as condition number when $M$ is a manifold, was introduced by Federer in 1959. The reach is a central concept in geometric measure theory, set estimation, manifold learning, among others areas. We introduce a universally consistent estimate of the reach, just assuming that the reach is positive. Under an additional assumption we provide rates of convergence. We also show that it is not possible to determine, based on a finite sample, if the reach of the support of a density is zero or not. We provide a small simulation study and a bias correction method for the case when $M$ is a manifold.

math.ST

Statistical analysis of measures of non-convexity

Several measures of non-convexity (departures from convexity) have been introduced in the literature, both for sets and functions. Some of them are of geometric nature, while others are more of topological nature. We address the statistical analysis of some of these measures of non-convexity of a set $S$, by dealing with their estimation based on a sample of points in $S$. We introduce also a new measure of non-convexity. We discuss briefly about these different notions of non-convexity, prove consistency and find the asymptotic distribution for the proposed estimators. We also consider the practical implementation of these estimators and illustrate their applicability to a real data example.

math.ST

Application of the Cramér-Wold theorem to testing for invariance under group actions

We address the problem of testing for the invariance of a probability measure under the action of a group of linear transformations. We propose a procedure based on consideration of one-dimensional projections, justified using a variant of the Cramér-Wold theorem. Our test procedure is powerful, computationally efficient, and dimension-independent, extending even to the case of infinite-dimensional spaces (multivariate functional data). It includes, as special cases, tests for exchangeability and sign-invariant exchangeability. We compare our procedure with some previous proposals in these cases, in a small simulation study. The paper concludes with two real-data examples.

math.ST

Level set and density estimation on manifolds

We tackle the problem of the estimation of the level sets L_f(λ) of the density f of a random vector X supported on a smooth manifold M\subsetR^d , from an iid sample of X. To do that we introduce a kernel-based estimator f^n,h , which is a slightly modified version of the one proposed in [45], and proves its a.s. uniform convergence to f . Then, we propose two estimators of L f (λ), the first one is a plug-in: L f^n,h (λ), which is proven to be a.s. consistent in Hausdorff distance and distance in measure, if L f(λ) does not meet the boundary of M . While the second one assumes that L f(λ) is r-convex, and is estimated by means of the r-convex hull of L f^n,h(λ). The performance of our proposal is illustrated through some simulated examples. In a real data example we analyze the intensity and direction of strong and moderate winds.

math.ST

Level sets of depth measures in abstract spaces

The lens depth of a point has been recently extended to general metric spaces, which is not the case for most depths. It is defined as the probability of being included in the intersection of two random balls centred at two random points X and Y, with the same radius d(X, Y). We study the consistency in Hausdorff and measure distance, of the level sets of the empirical lens depth, based on an iid sample on a general metric space. We also prove that the boundary of the empirical level sets are consistent estimators of their population counterparts, and analyze two real-life examples

math.ST

Sensitivity analysis in general metric spaces

In this paper, we introduce new indices adapted to outputs valued in general metric spaces. This new class of indices encompasses the classical ones; in particular, the so-called Sobol indices and the Cram{é}r-von-Mises indices. Furthermore, we provide asymptotically Gaussian estimators of these indices based on U-statistics. Surprisingly, we prove the asymp-totic normality straightforwardly. Finally, we illustrate this new procedure on a toy model and on two real-data examples.

math.ST

Weighted lens depth: Some applications to supervised classification

Starting with Tukey's pioneering work in the 1970's, the notion of depth in statistics has been widely extended especially in the last decade. These extensions include high dimensional data, functional data, and manifold-valued data. In particular, in the learning paradigm, the depth-depth method has become a useful technique. In this paper we extend the notion of lens depth to the case of data in metric spaces, and prove its main properties, with particular emphasis on the case of Riemannian manifolds, where we extend the concept of lens depth in such a way that it takes into account non-convex structures on the data distribution. Next we illustrate our results with some simulation results and also in some interesting real datasets, including pattern recognition in phylogenetic trees using the depth--depth approach.

math.ST