SearcharxivSearch

arXiv subjects

Marcela Svarc

Publications and source records attributed to Marcela Svarc.

10 recordsLinked to original sources

Mixed Gaussian Projections for Two-Sample Testing of Functional Data

Random-projection tests for functional data depend on the probability law used to generate projection directions. A measure has to be selected to generate the random directions in which the data is projected. In $L^2$, probability measures defined by their moments can be discriminated using Gaussian measures. The covariance operator of the Gaussian projection law determines which regions and structures of the functional space receive appreciable probability, and consequently affects finite-sample power. We show that mixtures of non-degenerate Gaussian measures preserve the almost-sure separation property and induce a metric between functional probability laws. A concentration argument further makes explicit that the power of projection tests is governed by the integrated projected distance associated with the chosen law. As a concrete construction, we combine Haar and Fourier Gaussian components, which emphasize localized and oscillatory departures, respectively. The resulting permutation test retains consistency, while component labels and upper-tail separation scores provide a descriptive indication of the geometry of the detected discrepancy. Simulations illustrate the specialization of the two components and the robustness of their mixture, and an ECG5000 application identifies a predominantly localized difference between two classes. We also compare with a pooled functional principal component analysis covariance operator. This finite-rank, data-adaptive projection geometry preserves permutation validity through its label-invariant construction and improves power in several simulated settings, particularly under a change in covariance structure.

stat.ME

The brain uses renewal points to model random sequences of stimuli

It has been classically conjectured that the brain assigns probabilistic models to sequences of stimuli. An important issue associated with this conjecture is the identification of the classes of models used by the brain to perform this task. We address this issue by using a new clustering procedure for sets of electroencephalographic (EEG) data recorded from participants exposed to a sequence of auditory stimuli generated by a stochastic chain. This clustering procedure indicates that the brain uses renewal points in the stochastic sequence of auditory stimuli in order to build a model.

q-bio.NC

Clustering Sets of Functional Data by Similarity in Law

We introduce a new clustering method for the classification of functional data sets by their probabilistic law, that is, a procedure that aims to assign data sets to the same cluster if and only if the data were generated with the same underlying distribution. This method has the nice virtue of being non-supervised and non-parametric, allowing for exploratory investigation with few assumptions about the data. Rigorous finite bounds on the classification error are given along with an objective heuristic that consistently selects the best partition in a data-driven manner. Simulated data has been clustered with this procedure to show the performance of the method with different parametric model classes of functional data.

stat.ME

A local depth measure for general data

We introduce the Integrated Dual Local Depth which is a local depth measure for data in a Banach space based on the use of one-dimensional projections. The properties of a depth measure are analyzed under this setting and a proper definition of local symmetry is given. Moreover, strong consistency results for the local depth and also for the local depth regions are attained. Finally, applications to descriptive data analysis and classification are analyzed, making the special focus on multivariate functional data, where we obtain very promising results.

stat.ME

Sequential Clustering for Functional Data

This paper presents SeqClusFD, a top-down sequential clustering method for functional data. The clustering algorithm extracts the splitting information either from trajectories, first or second derivatives. Initial partition is based on gap statistic that provides local information to identify the instant with more clustering evidence in trajectories or derivatives. Then functional boxplots allow reconsidering overall allocation and each observation is finally assigned to the cluster where it spends most of the time within whiskers. These local and global searches are repeated recursively until there is no evidence of clustering at any time on trajectories or first and second derivatives. SeqClusFD simultaneously estimates the number of groups and provides data allocation. It also provides valuable information about the most important features that determine cluster structure. Computational aspects have been analyzed and the new method is tested on synthetic and real data sets.

stat.ME

Feature Selection for Functional Data

In this paper we address the problem of feature selection when the data is functional, we study several statistical procedures including classification, regression and principal components. One advantage of the blinding procedure is that it is very flexible since the features are defined by a set of functions, relevant to the problem being studied, proposed by the user. Our method is consistent under a set of quite general assumptions, and produces good results with the real data examples that we analyze.

stat.ME

Clustering using Unsupervised Binary Trees: CUBT

We herein introduce a new method of interpretable clustering that uses unsupervised binary trees. It is a three-stage procedure, the first stage of which entails a series of recursive binary splits to reduce the heterogeneity of the data within the new subsamples. During the second stage (pruning), consideration is given to whether adjacent nodes can be aggregated. Finally, during the third stage (joining), similar clusters are joined together, even if they do not share the same parent originally. Consistency results are obtained, and the procedure is used on simulated and real data sets.

stat.ME

Interpretable Clustering using Unsupervised Binary Trees

We herein introduce a new method of interpretable clustering that uses unsupervised binary trees. It is a three-stage procedure, the first stage of which entails a series of recursive binary splits to reduce the heterogeneity of the data within the new subsamples. During the second stage (pruning), consideration is given to whether adjacent nodes can be aggregated. Finally, during the third stage (joining), similar clusters are joined together, even if they do not descend from the same node originally. Consistency results are obtained, and the procedure is used on simulated and real data sets.

stat.ME

Resistant estimates for high dimensional and functional data based on random projections

We herein propose a new robust estimation method based on random projections that is adaptive and, automatically produces a robust estimate, while enabling easy computations for high or infinite dimensional data. Under some restricted contamination models, the procedure is robust and attains full efficiency. We tested the method using both simulated and real data.

stat.ME

Selection of variables for cluster analysis and classification rules

In this paper we introduce two procedures for variable selection in cluster analysis and classification rules. One is mainly oriented to detect the noisy non-informative variables, while the other deals also with multicolinearity. A forward-backward algorithm is also proposed to make feasible these procedures in large data sets. A small simulation is performed and some real data examples are analyzed.

math.ST