SearcharxivSearch

arXiv subjects

Alicia Nieto-Reyes

Publications and source records attributed to Alicia Nieto-Reyes.

15 recordsLinked to original sources

seMCD: Sequentially implemented Monte Carlo depth computation with statistical guarantees

Statistical depth functions provide center-outward orderings in spaces of dimension larger than one, where a natural ordering does not exist. The numerical evaluation of such depth functions can be computationally prohibitive, even for relatively low dimensions. We present a novel sequentially implemented Monte Carlo methodology for the computation of, theoretical and empirical, depth functions and related quantities (seMCD), that outputs an interval, a so-called seMCD-bucket, to which the quantity of interest belongs with a high probability prespecified by the user. For specific classes of depth functions, we adapt algorithms from sequential testing, providing finite-sample guarantees. For depth functions dependent on unknown distributions, we offer asymptotic guarantees using non-parametric statistical methods. In contrast to plain-vanilla Monte Carlo methodology the number of samples required in the algorithm is random but typically much smaller than standard choices suggested in the literature. The seMCD method can be applied to various depth functions, covering multivariate and functional spaces. We demonstrate the efficiency and reliability of our approach through empirical studies, highlighting its applicability in outlier or anomaly detection, classification, and depth region computation. In conclusion, the seMCD-algorithm can achieve accurate depth approximations with few Monte Carlo samples while maintaining rigorous statistical guarantees.

stat.ME

Symmetrisation of a class of two-sample tests by mutually considering depth ranks including functional spaces

Statistical depth functions provide measures of the outlyingness, or centrality, of the elements of a space with respect to a distribution. It is a nonparametric concept applicable to spaces of any dimension, for instance, multivariate and functional. Liu and Singh (1993) presented a multivariate two-sample test based on depth-ranks. We dedicate this paper to improving the power of the associated test statistic and incorporating its applicability to functional data. In doing so, we obtain a more natural test statistic that is symmetric in both samples. We derive the null asymptotic of the proposed test statistic, also proving the validity of the testing procedure for functional data. Finally, the finite sample performance of the test for functional data is illustrated by means of a simulation study and a real data analysis on annual temperature curves of ocean drifters is executed.

math.ST

Generalization of the simplicial depth: no vanishment outside the convex hull of the distribution support

The simplicial depth, like other relevant multivariate statistical data depth functions, vanishes right outside the convex hull of the support of the distribution with respect to which the depth is computed. This is problematic when it is required to differentiate among points outside the convex hull of the distribution support, with respect to which the depth is computed, based on their depth values. We provide the first proposal for simplicial depth which do not vanish right outside the convex hull of the distribution. The properties of the proposal and of the corresponding estimator are studied theoretically and by means of Monte Carlo simulations and analysis of datasets.

math.ST

Statistical depth and support medians for fuzzy data

Statistical depth functions order the elements of a space with respect to their centrality in a probability distribution or dataset. Since many depth functions are maximized in the real line by the median, they provide a natural approach to defining median-like location estimators for more general types of data (in our case, fuzzy data). We analyze the relationships between depth-based medians, medians based on the support function, and some notions of a median for fuzzy data in the literature. We take advantage of specific depth functions for fuzzy data defined in our former papers: adaptations of Tukey depth, simplicial depth, $L^1$-depth and projection depth.

math.ST

Projection depth and $L^r$-type depths for fuzzy random variables

Statistical depth functions are a standard tool in nonparametric statistics to extend order-based univariate methods to the multivariate setting. Since there is no universally accepted total order for fuzzy data (even in the univariate case) and there is a lack of parametric models, a fuzzy extension of depth-based methods is very interesting. In this paper, we adapt projection depth and $L^{r}$-type depth to the fuzzy setting, studying their properties and illustrating their behaviour with a real data example.

math.ST

Analytical and statistical properties of local depth functions motivated by clustering applications

Local general depth ($LGD$) functions are used for describing the local geometric features and mode(s) in multivariate distributions. In this paper, we undertake a rigorous systematic study of $LGD$ and establish several analytical and statistical properties. First, we show that, when the underlying probability distribution is absolutely continuous with density $f(\cdot)$, the scaled version of $LGD$ (referred to as $τ$-approximation) converges, uniformly and in $L^d(\mathbb{R}^p)$ to $f(\cdot)$ when $τ$ converges to zero. Second, we establish that, as the sample size diverges to infinity the centered and scaled sample $LGD$ converge in distribution to a centered Gaussian process uniformly in the space of bounded functions on $\mathcal{H}_G$, a class of functions yielding $LGD$. Third, using the sample version of the $τ$-approximation ($S τA$) and the gradient system analysis, we develop a new clustering algorithm. The validity of this algorithm requires several results concerning the uniform finite difference approximation of the gradient system associated with $S τA$. For this reason, we establish \emph{Bernstein}-type inequality for deviations between the centered and scaled sample $LGD$, which is also of independent interest. Finally, invoking the above results, we establish consistency of the clustering algorithm. Applications of the proposed methods to mode estimation and upper level set estimation are also provided. Finite sample performance of the methodology are evaluated using numerical experiments and data analysis.

math.ST

On the Non-Gaussianity of Sea Surface Elevations

The sea surface elevations are generally stated as Gaussian processes in the literature. To show the inaccuracy of this statement, an empirical study of the buoys in the US coast at a random day is performed, which results in rejecting the null hypothesis of Gaussianity in over 80$\%$ of the cases. The analysis pursued relates to a recent one by the author in which the heights of sea waves are proved to be non-Gaussian. It is similar in that the Gaussianity of the process is studied as a whole and not just of its one-dimensional marginal, as it is common in the literature. It differs, however, in that the analysis of the sea surface elevations is harder from a statistical point of view, as the one-dimensional marginals are commonly Gaussian, which is observed throughout the study.

stat.AP

Properties of statistical depth with respect to compact convex random sets. The Tukey depth

We study a statistical data depth with respect to compact convex random sets which is consistent with the multivariate Tukey depth and the Tukey depth for fuzzy sets. In doing so, we provide a series of properties for statistical data depth with respect to compact convex random sets. These properties are an adaptation of properties that constitute the axiomatic notions of multivariate, functional and fuzzy depth functions and other well-known properties of depth.

math.ST

Simplicial depths for fuzzy random variables

The recently defined concept of a statistical depth function for fuzzy sets provides a theoretical framework for ordering fuzzy sets with respect to the distribution of a fuzzy random variable. One of the most used and studied statistical depth function for multivariate data is simplicial depth, based on multivariate simplices. We introduce a notion of pseudosimplices generated by fuzzy sets and propose three plausible generalizations of simplicial depth to fuzzy sets. Their theoretical properties are analyzed and the behavior of the proposals illustrated through a study of both synthetic and real data.

stat.ME

Statistical Depth for Big Functional Data with Application to Neuroimaging

Functional depth is the functional data analysis technique that orders a functional data set. Unlike the case of data on the real line, defining this order is non-trivial, and particularly, with functional data, there are a number of properties that any depth should satisfy. We propose a new depth which both satisfies the properties required of a functional depth but also one which can be used in the case where there are a very large number of functional observations or in the case where the observations are functions of several continuous variables (such as images, for example). We give theoretical justification for our choice, and evaluate our proposed depth through simulation. We finally apply the proposed depth to the problem of yielding a completely non-parametric deconvolution of Positron Emission Tomography (PET) data for a very large number of curves across the image, as well as to the problem of finding a representative subject from a set of PET scans.

stat.ME

Statistical Depth based Normalization and Outlier Detection of Gene Expression Data

Normalization and outlier detection belong to the preprocessing of gene expression data. We propose a natural normalization procedure based on statistical data depth which normalizes to the distribution of gene expressions of the most representative gene expression of the group. This differ from the standard method of quantile normalization, based on the coordinate-wise median array that lacks of the well-known properties of the one-dimensional median. The statistical data depth maintains those good properties. Gene expression data are known for containing outliers. Although detecting outlier genes in a given gene expression dataset has been broadly studied, these methodologies do not apply for detecting outlier samples, given the difficulties posed by the high dimensionality but low sample size structure of the data. The standard procedures used for detecting outlier samples are visual and based on dimension reduction techniques; instances are multidimensional scaling and spectral map plots. For detecting outlier genes in a given gene expression dataset, we propose an analytical procedure and based on the Tukey's concept of outlier and the notion of statistical depth, as previous methodologies lead to unassertive and wrongful outliers. We reveal the outliers of four datasets; as a necessary step for further research.

stat.ME

Statistical depth in abstract metric spaces

The concept of depth has proved very important for multivariate and functional data analysis, as it essentially acts as a surrogate for the notion a ranking of observations which is absent in more than one dimension. Motivated by the rapid development of technology, in particular the advent of `Big Data', we extend here that concept to general metric spaces, propose a natural depth measure and explore its properties as a statistical depth function. Working in a general metric space allows the depth to be tailored to the data at hand and to the ultimate goal of the analysis, a very desirable property given the polymorphic nature of modern data sets. This flexibility is thoroughly illustrated by several real data analyses.

stat.ME

An R package for Normality in Stationary Processes

Normality is the main assumption for analyzing dependent data in several time series models, and tests of normality have been widely studied in the literature, however, the implementations of these tests are limited. The \textbf{nortsTest} package performs the tests of \textit{Lobato and Velasco, Epps, Psaradakis and Vavra} and \textit{random projection} for normality of stationary processes. In addition, the package offers visual diagnostics for checking stationarity and normality assumptions for the most used time series models in several \R packages. The aim of this work is to show the functionality of the package, presenting each test performance with simulated examples, and the package utility for model diagnostic in time series analysis.

stat.CO

A topologically valid definition of depth for functional data

The main focus of this work is on providing a formal definition of statistical depth for functional data on the basis of six properties, recognising topological features such as continuity, smoothness and contiguity. Amongst our depth defining properties is one that addresses the delicate challenge of inherent partial observability of functional data, with fulfilment giving rise to a minimal guarantee on the performance of the empirical depth beyond the idealised and practically infeasible case of full observability. As an incidental product, functional depths satisfying our definition achieve a robustness that is commonly ascribed to depth, despite the absence of a formal guarantee in the multivariate definition of depth. We demonstrate the fulfilment or otherwise of our properties for six widely used functional depth proposals, thereby providing a systematic basis for selection of a depth function.

math.ST

Linear free divisors and the global logarithmic comparison theorem

A complex hypersurface D in complex affine n-space C^n is a linear free divisor (LFD) if its module of logarithmic vector fields has a global basis of linear vector fields. We classify all LFDs for n at most 4. Analogous to Grothendieck's comparison theorem, we say that the global logarithmic comparison theorem (GLCT) holds for D if the complex of global logarithmic differential forms computes the complex cohomology of the complement of D in C^n. We develop a general criterion for the GLCT for LFDs and prove that it is fulfilled whenever the Lie algebra of linear logarithmic vector fields is reductive. For n at most 4, we show that the GLCT holds for all LFDs. We show that LFDs arising naturally as discriminants in quiver representation spaces (of real Schur roots) fulfill the GLCT. As a by-product we obtain a simplified proof of a theorem of V. Kac on the number of irreducible components of such discriminants.

math.AG