SearcharxivSearch

arXiv subjects

Carlo Sguera

Publications and source records attributed to Carlo Sguera.

5 recordsLinked to original sources

Detecting and Classifying Outliers in Big Functional Data

We propose two new outlier detection methods, for identifying and classifying different types of outliers in (big) functional data sets. The proposed methods are based on an existing method called Massive Unsupervised Outlier Detection (MUOD). MUOD detects and classifies outliers by computing for each curve, three indices, all based on the concept of linear regression and correlation, which measure outlyingness in terms of shape, magnitude and amplitude, relative to the other curves in the data. 'Semifast-MUOD', the first method, uses a sample of the observations in computing the indices, while 'Fast-MUOD', the second method, uses the point-wise or $L_1$ median in computing the indices. The classical boxplot is used to separate the indices of the outliers from those of the typical observations. Performance evaluation of the proposed methods using simulated data show significant improvements compared to MUOD, both in outlier detection and computational time. We show that Fast-MUOD is especially well suited to handling big and dense functional datasets with very small computational time compared to other methods. Further comparisons with some recent outlier detection methods for functional data also show superior or comparable outlier detection accuracy of the proposed methods. We apply the proposed methods on weather, population growth, and video data.

stat.ME

A notion of depth for sparse functional data

Data depth is a well-known and useful nonparametric tool for analyzing functional data. It provides a novel way of ranking a sample of curves from the center outwards and defining robust statistics, such as the median or trimmed means. It has also been used as a building block for functional outlier detection methods and classification. Several notions of depth for functional data were introduced in the literature in the last few decades. These functional depths can only be directly applied to samples of curves measured on a fine and common grid. In practice, this is not always the case, and curves are often observed at sparse and subject dependent grids. In these scenarios the usual approach consists in estimating the trajectories on a common dense grid, and using the estimates in the depth analysis. This approach ignores the uncertainty associated with the curves estimation step. Our goal is to extend the notion of depth so that it takes into account this uncertainty. Using both functional estimates and their associated confidence intervals, we propose a new method that allows the curve estimation uncertainty to be incorporated into the depth analysis. We describe the new approach using the modified band depth although any other functional depth could be used. The performance of the proposed methodology is illustrated using simulated curves in different settings where we control the degree of sparsity. Also a real data set consisting of female medflies egg-laying trajectories is considered. The results show the benefits of using uncertainty when computing depth for sparse functional data.

stat.ME

An empirical comparison of global and local functional depths

A functional data depth provides a center-outward ordering criterion which allows the definition of measures such as median, trimmed means, central regions or ranks in a functional framework. A functional data depth can be global or local. With global depths, the degree of centrality of a curve $x$ depends equally on the rest of the sample observations, while with local depths, the contribution of each observation in defining the degree of centrality of $x$ decreases as the distance from $x$ increases. We empirically compare the global and the local approaches to the functional depth problem focusing on three global and two local functional depths. First, we consider two real data sets and show that global and local depths may provide different insights. Second, we use simulated data to show when we should expect differences between a global and a local approach to the functional depth problem.

stat.ME

Functional outlier detection by a local depth with application to NOx levels

This paper proposes methods to detect outliers in functional data sets and the task of identifying atypical curves is carried out using the recently proposed kernelized functional spatial depth (KFSD). KFSD is a local depth that can be used to order the curves of a sample from the most to the least central, and since outliers are usually among the least central curves, we present a probabilistic result which allows to select a threshold value for KFSD such that curves with depth values lower than the threshold are detected as outliers. Based on this result, we propose three new outlier detection procedures. The results of a simulation study show that our proposals generally outperform a battery of competitors. We apply our procedures to a real data set consisting in daily curves of emission levels of nitrogen oxides (NOx) since it is of interest to identify abnormal NOx levels to take necessary environmental political actions.

stat.ME

Spatial Depth-Based Classification for Functional Data

We enlarge the number of available functional depths by introducing the kernelized functional spatial depth (KFSD). KFSD is a local-oriented and kernel-based version of the recently proposed functional spatial depth (FSD) that may be useful for studying functional samples that require an analysis at a local level. In addition, we consider supervised functional classification problems, focusing on cases in which the differences between groups are not extremely clear-cut or the data may contain outlying curves. We perform classification by means of some available robust methods that involve the use of a given functional depth, including FSD and KFSD, among others. We use the functional \textit{k}-nearest neighbor classifier as a benchmark procedure. The results of a simulation study indicate that the KFSD-based classification approach leads to good results. Finally, we consider two real classification problems, obtaining results that are consistent with the findings observed with simulated curves.

stat.ME