SearcharxivSearch

arXiv subjects

Antonio Cuevas

Publications and source records attributed to Antonio Cuevas.

At least 19 recordsLinked to original sources

On consistent estimation of dimension values

The problem of estimating, from a random sample of points, the dimension of a compact subset $S$ of the Euclidean space is considered. The emphasis is put on consistency results in the statistical sense. That is, statements of convergence to the true dimension value when the sample size grows to infinity. Among the many available definitions of dimension, we have focused (on the grounds of its statistical tractability) on three notions: the Minkowski dimension, the correlation dimension and the, perhaps less popular, concept of pointwise dimension. We prove the statistical consistency of some natural estimators of these quantities. Our proofs partially rely on the use of an instrumental estimator formulated in terms of the empirical volume function $V_n(r)$, defined as the Lebesgue measure of the set of points whose distance to the sample is at most $r$. In particular, we explore the case in which the true volume function $V(r)$ of the target set $S$ is a polynomial on some interval starting at zero. An empirical study is also included. Our study aims to provide some theoretical support, and some practical insights, for the problem of deciding whether or not the set $S$ has a dimension smaller than that of the ambient space. This is a major statistical motivation of the dimension studies, in connection with the so-called ``Manifold Hypothesis''.

math.ST

A Bayesian approach to functional regression: theory and computation

We propose a novel Bayesian methodology for inference in functional linear and logistic regression models based on the theory of reproducing kernel Hilbert spaces (RKHS's). We introduce general models that build upon the RKHS generated by the covariance function of the underlying stochastic process, and whose formulation includes as particular cases all finite-dimensional models based on linear combinations of marginals of the process, which can collectively be seen as a dense subspace made of simple approximations. By imposing a suitable prior distribution on this dense functional space we can perform data-driven inference via standard Bayes methodology, estimating the posterior distribution through reversible jump Markov chain Monte Carlo methods. In this context, our contribution is two-fold. First, we derive theoretical results that guarantee strong posterior consistency and contraction at an optimal rate under mild conditions. Second, we show that several prediction strategies stemming from our Bayesian procedure are competitive against other usual alternatives in both simulations and real data sets, including a Bayesian-motivated variable selection method.

stat.ME

On uniqueness of the set of k-means

We provide necessary and sufficient conditions for the uniqueness of the k-means set of a probability distribution. This uniqueness problem is related to the choice of k: depending on the underlying distribution, some values of this parameter could lead to multiple sets of k-means, which hampers the interpretation of the results and/or the stability of the algorithms. We give a general assessment on consistency of the empirical k-means adapted to the setting of non-uniqueness and determine the asymptotic distribution of the within cluster sum of squares (WCSS). We also provide statistical characterizations of k-means uniqueness in terms of the asymptotic behavior of the empirical WCSS. As a consequence, we derive a bootstrap test for uniqueness of the set of k-means. The results are illustrated with examples of different types of non-uniqueness and we check by simulations the performance of the proposed methodology.

math.ST

A uniform kernel trick for high-dimensional two-sample problems

We use a suitable version of the so-called "kernel trick" to devise two-sample (homogeneity) tests, especially focussed on high-dimensional and functional data. Our proposal entails a simplification related to the important practical problem of selecting an appropriate kernel function. Specifically, we apply a uniform variant of the kernel trick which involves the supremum within a class of kernel-based distances. We obtain the asymptotic distribution (under the null and alternative hypotheses) of the test statistic. The proofs rely on empirical processes theory, combined with the delta method and Hadamard (directional) differentiability techniques, and functional Karhunen-Loève-type expansions of the underlying processes. This methodology has some advantages over other standard approaches in the literature. We also give some experimental insight into the performance of our proposal compared to the original kernel-based approach \cite{Gretton2007} and the test based on energy distances \cite{Szekely-Rizzo-2017}.

math.ST

On the notion of polynomial reach: a statistical application

The volume function V(t) of a compact set S\in R^d is just the Lebesgue measure of the set of points within a distance to S not larger than t. According to some classical results in geometric measure theory, the volume function turns out to be a polynomial, at least in a finite interval, under a quite intuitive, easy to interpret, sufficient condition (called ``positive reach'') which can be seen as an extension of the notion of convexity. However, many other simple sets, not fulfilling the positive reach condition, have also a polynomial volume function. To our knowledge, there is no general, simple geometric description of such sets. Still, the polynomial character of $V(t)$ has some relevant consequences since the polynomial coefficients carry some useful geometric information. In particular, the constant term is the volume of S and the first order coefficient is the boundary measure (in Minkowski's sense). This paper is focused on sets whose volume function is polynomial on some interval starting at zero, whose length (that we call ``polynomial reach'') might be unknown. Our main goal is to approximate such polynomial reach by statistical means, using only a large enough random sample of points inside S. The practical motivation is simple: when the value of the polynomial reach , or rather a lower bound for it, is approximately known, the polynomial coefficients can be estimated from the sample points by using standard methods in polynomial approximation. As a result, we get a quite general method to estimate the volume and boundary measure of the set, relying only on an inner sample of points and not requiring the use any smoothing parameter. This paper explores the theoretical and practical aspects of this idea.

math.ST

On functional logistic regression: some conceptual issues

The main ideas behind the classical multivariate logistic regression model make sense when translated to the functional setting, where the explanatory variable $X$ is a function and the response $Y$ is binary. However, some important technical issues appear (or are aggravated with respect to those of the multivariate case) due to the functional nature of the explanatory variable. First, the mere definition of the model can be questioned: while most approaches so far proposed rely on the $L_2$-based model, we suggest an alternative (in some sense, more general) approach, based on the theory of Reproducing Kernel Hilbert Spaces (RKHS). The validity conditions of such RKHS-based model, as well as its relation with the $L_2$-based one are investigated and made explicit in two formal results. Some relevant particular cases are considered as well. Second we show that, under very general conditions, the maximum likelihood (ML) of the logistic model parameters fail to exist in the functional case. Third, on a more positive side, we suggest an RKHS-based restricted version of the ML estimator. This is a methodological paper, aimed at a better understanding of the functional logistic model, rather than focusing on numerical and practical issues.

math.ST

On a general definition of the functional linear model

A general formulation of the linear model with functional (random) explanatory variable $X = X(t), t \in T$ , and scalar response Y is proposed. It includes the standard functional linear model, based on the inner product in the space $L^2[0,1]$, as a particular case. It also includes all models in which Y is assumed to be (up to an additive noise) a linear combination of a finite or countable collections of marginal variables X(t_j), with $t_j\in T$ or a linear combination of a finite number of linear projections of X. This general formulation can be interpreted in terms of the RKHS space generated by the covariance function of the process X(t). Some consistency results are proved. A few experimental results are given in order to show the practical interest of considering, in a unified framework, linear models based on a finite number of marginals $X(t_j)$ of the process $X(t)$.

math.ST

Set Estimation Under Biconvexity Restrictions

A set in the Euclidean plane is said to be biconvex if, for some angle $θ\in[0,π/2)$, all its sections along straight lines with inclination angles $θ$ and $θ+π/2$ are convex sets (i.e, empty sets or segments). Biconvexity is a natural notion with some useful applications in optimization theory. It has also be independently used, under the name of "rectilinear convexity", in computational geometry. We are concerned here with the problem of asymptotically reconstructing (or estimating) a biconvex set $S$ from a random sample of points drawn on $S$. By analogy with the classical convex case, one would like to define the "biconvex hull" of the sample points as a natural estimator for $S$. However, as previously pointed out by several authors, the notion of "hull" for a given set $A$ (understood as the "minimal" set including $A$ and having the required property) has no obvious, useful translation to the biconvex case. This is in sharp contrast with the well-known elementary definition of convex hull. Thus, we have selected the most commonly accepted notion of "biconvex hull" (often called "rectilinear convex hull"): we first provide additional motivations for this definition, proving some useful relations with other convexity-related notions. Then, we prove some results concerning the consistent approximation of a biconvex set $S$ and and the corresponding biconvex hull. An analogous result is also provided for the boundaries. A method to approximate, from a sample of points on $S$, the biconvexity angle $θ$ is also given.

math.ST

Directional differentiability for supremum-type functionals: statistical applications

We show that various functionals related to the supremum of a real function defined on an arbitrary set or a measure space are Hadamard directionally differentiable. We specifically consider the supremum norm, the supremum, the infimum, and the amplitude of a function. The (usually non-linear) derivatives of these maps adopt simple expressions under suitable assumptions on the underlying space. As an application, we improve and extend to the multidimensional case the results in \cite{Raghavachari} regarding the limiting distributions of Kolmogorov-Smirnov type statistics under the alternative hypothesis. Similar results are obtained for analogous statistics associated with copulas. We additionally solve an open problem about the Berk-Jones statistic proposed by \cite{Jager-Wellner-2004}. Finally, the asymptotic distribution of maximum mean discrepancies over Donsker classes of functions is derived.

math.ST

On Mahalanobis distance in functional settings

Mahalanobis distance is a classical tool in multivariate analysis. We suggest here an extension of this concept to the case of functional data. More precisely, the proposed definition concerns those statistical problems where the sample data are real functions defined on a compact interval of the real line. The obvious difficulty for such a functional extension is the non-invertibility of the covariance operator in infinite-dimensional cases. Unlike other recent proposals, our definition is suggested and motivated in terms of the Reproducing Kernel Hilbert Space (RKHS) associated with the stochastic process that generates the data. The proposed distance is a true metric; it depends on a unique real smoothing parameter which is fully motivated in RKHS terms. Moreover, it shares some properties of its finite dimensional counterpart: it is invariant under isometries, it can be consistently estimated from the data and its sampling distribution is known under Gaussian models. An empirical study for two statistical applications, outliers detection and binary classification, is included. The obtained results are quite competitive when compared to other recent proposals of the literature.

stat.ME

Detection of low dimensionality and data denoising via set estimation techniques

This work is closely related to the theories of set estimation and manifold estimation. Our object of interest is a, possibly lower-dimensional, compact set $S \subset {\mathbb R}^d$. The general aim is to identify (via stochastic procedures) some qualitative or quantitative features of $S$, of geometric or topological character. The available information is just a random sample of points drawn on $S$. The term "to identify" means here to achieve a correct answer almost surely (a.s.) when the sample size tends to infinity. More specifically the paper aims at giving some partial answers to the following questions: is $S$ full dimensional? Is $S$ "close to a lower dimensional set" $\mathcal{M}$? If so, can we estimate $\mathcal{M}$ or some functionals of $\mathcal{M}$ (in particular, the Minkowski content of $\mathcal{M}$)? As an important auxiliary tool in the answers of these questions, a denoising procedure is proposed in order to partially remove the noise in the original data. The theoretical results are complemented with some simulations and graphical illustrations.

math.ST

An RKHS model for variable selection in functional regression

A mathematical model for variable selection in functional regression models with scalar response is proposed. By "variable selection" we mean a procedure to replace the whole trajectories of the functional explanatory variables with their values at a finite number of carefully selected instants (or "impact points"). The basic idea of our approach is to use the Reproducing Kernel Hilbert Space (RKHS) associated with the underlying process, instead of the more usual L2[0,1] space, in the definition of the linear model. This turns out to be especially suitable for variable selection purposes, since the finite-dimensional linear model based on the selected "impact points" can be seen as a particular case of the RKHS-based linear functional model. In this framework, we address the consistent estimation of the optimal design of impact points and we check, via simulations and real data examples, the performance of the proposed method.

stat.ME

On the use of reproducing kernel Hilbert spaces in functional classification

The Hájek-Feldman dichotomy establishes that two Gaussian measures are either mutually absolutely continuous with respect to each other (and hence there is a Radon-Nikodym density for each measure with respect to the other one) or mutually singular. Unlike the case of finite dimensional Gaussian measures, there are non-trivial examples of both situations when dealing with Gaussian stochastic processes. This paper provides: (a) Explicit expressions for the optimal (Bayes) rule and the minimal classification error probability in several relevant problems of supervised binary classification of mutually absolutely continuous Gaussian processes. The approach relies on some classical results in the theory of Reproducing Kernel Hilbert Spaces (RKHS). (b) An interpretation, in terms of mutual singularity, for the "near perfect classification" phenomenon described by Delaigle and Hall (2012). We show that the asymptotically optimal rule proposed by these authors can be identified with the sequence of optimal rules for an approximating sequence of classification problems in the absolutely continuous case. (c) A new model-based method for variable selection in binary classification problems, which arises in a very natural way from the explicit knowledge of the RN-derivatives and the underlying RKHS structure. Different classifiers might be used from the selected variables. In particular, the classical, linear finite-dimensional Fisher rule turns out to be consistent under some standard conditions on the underlying functional model.

stat.ME

On visual distances for spectrum-type functional data

A functional distance ${\mathbb H}$, based on the Hausdorff metric between the function hypographs, is proposed for the space ${\mathcal E}$ of non-negative real upper semicontinuous functions on a compact interval. The main goal of the paper is to show that the space $({\mathcal E},{\mathbb H})$ is particularly suitable in some statistical problems with functional data which involve functions with very wiggly graphs and narrow, sharp peaks. A typical example is given by spectrograms, either obtained by magnetic resonance or by mass spectrometry. On the theoretical side, we show that $({\mathcal E},{\mathbb H})$ is a complete, separable locally compact space and that the ${\mathbb H}$-convergence of a sequence of functions implies the convergence of the respective maximum values of these functions. The probabilistic and statistical implications of these results are discussed in particular, regarding the consistency of $k$-NN classifiers for supervised classification problems with functional data in ${\mathbb H}$. On the practical side, we provide the results of a small simulation study and check also the performance of our method in two real data problems of supervised classification involving mass spectra.

math.ST

The mRMR variable selection method: a comparative study for functional data

The use of variable selection methods is particularly appealing in statistical problems with functional data. The obvious general criterion for variable selection is to choose the `most representative' or `most relevant' variables. However, it is also clear that a purely relevance-oriented criterion could lead to select many redundant variables. The mRMR (minimum Redundance Maximum Relevance) procedure, proposed by Ding and Peng (2005) and Peng et al. (2005) is an algorithm to systematically perform variable selection, achieving a reasonable trade-off between relevance and redundancy. In its original form, this procedure is based on the use of the so-called mutual information criterion to assess relevance and redundancy. Keeping the focus on functional data problems, we propose here a modified version of the mRMR method, obtained by replacing the mutual information by the new association measure (called distance correlation) suggested by Székely et al. (2007). We have also performed an extensive simulation study, including 1600 functional experiments (100 functional models $\times$ 4 sample sizes $\times$ 4 classifiers) and three real-data examples aimed at comparing the different versions of the mRMR methodology. The results are quite conclusive in favor of the new proposed alternative.

stat.ME

Variable selection in functional data classification: a maxima-hunting proposal

Variable selection is considered in the setting of supervised binary classification with functional data $\{X(t),\ t\in[0,1]\}$. By "variable selection" we mean any dimension-reduction method which leads to replace the whole trajectory $\{X(t),\ t\in[0,1]\}$, with a low-dimensional vector $(X(t_1),\ldots,X(t_k))$ still keeping a similar classification error. Our proposal for variable selection is based on the idea of selecting the local maxima $(t_1,\ldots,t_k)$ of the function ${\mathcal V}_X^2(t)={\mathcal V}^2(X(t),Y)$, where ${\mathcal V}$ denotes the "distance covariance" association measure for random variables due to Székely, Rizzo and Bakirov (2007). This method provides a simple natural way to deal with the relevance vs. redundancy trade-off which typically appears in variable selection. This paper includes (a) Some theoretical motivation: a result of consistent estimation on the maxima of ${\mathcal V}_X^2$ is shown. We also show different theoretical models for the underlying process $X(t)$ under which the relevant information in concentrated in the maxima of ${\mathcal V}_X^2$. (b) An extensive empirical study, including about 400 simulated models and real data examples, aimed at comparing our variable selection method with other standard proposals for dimension reduction.

stat.ME

A geometrically motivated parametric model in manifold estimation,

The general aim of manifold estimation is reconstructing, by statistical methods, an $m$-dimensional compact manifold $S$ on ${\mathbb R}^d$ (with $m\leq d$) or estimating some relevant quantities related to the geometric properties of $S$. We will assume that the sample data are given by the distances to the $(d-1)$-dimensional manifold $S$ from points randomly chosen on a band surrounding $S$, with $d=2$ and $d=3$. The point in this paper is to show that, if $S$ belongs to a wide class of compact sets (which we call \it sets with polynomial volume\rm), the proposed statistical model leads to a relatively simple parametric formulation. In this setup, standard methodologies (method of moments, maximum likelihood) can be used to estimate some interesting geometric parameters, including curvatures and Euler characteristic. We will particularly focus on the estimation of the $(d-1)$-dimensional boundary measure (in Minkowski's sense) of $S$. It turns out, however, that the estimation problem is not straightforward since the standard estimators show a remarkably pathological behavior: while they are consistent and asymptotically normal, their expectations are infinite. The theoretical and practical consequences of this fact are discussed in some detail.

math.ST

On Poincaré cone property

A domain $S\subset{\mathbb{R}}^d$ is said to fulfill the Poincaré cone property if any point in the boundary of $S$ is the vertex of a (finite) cone which does not otherwise intersects the closure $\bar{S}$. For more than a century, this condition has played a relevant role in the theory of partial differential equations, as a shape assumption aimed to ensure the existence of a solution for the classical Dirichlet problem on $S$. In a completely different setting, this paper is devoted to analyze some statistical applications of the Poincaré cone property (when defined in a slightly stronger version). First, we show that this condition can be seen as a sort of generalized convexity: while it is considerably less restrictive than convexity, it still retains some ``convex flavour.'' In particular, when imposed to a probability support $S$, this property allows the estimation of $S$ from a random sample of points, using the ``hull principle'' much in the same way as a convex support is estimated using the convex hull of the sample points. The statistical properties of such hull estimator (consistency, convergence rates, boundary estimation) are considered in detail. Second, it is shown that the class of sets fulfilling the Poincaré property is a $P$-Glivenko-Cantelli class for any absolutely continuous distribution $P$ on $\mathbb{R}^d$. This has some independent interest in the theory of empirical processes, since it extends the classical analogous result, established for convex sets, to a much larger class. Third, an algorithm to approximate the cone-convex hull of a finite sample of points is proposed and some practical illustrations are given.

math.ST