SearcharxivSearch

arXiv subjects

R. Fraiman

Publications and source records attributed to R. Fraiman.

3 recordsLinked to original sources

Sensitivity indices for output on a Riemannian manifold

In the context of computer code experiments, sensitivity analysis of a complicated input-output system is often performed by ranking the so-called Sobol indices. One reason of the popularity of Sobol's approach relies on the simplicity of the statistical estimation of these indices using the so-called Pick and Freeze method. In this work we propose and study sensitivity indices for the case where the output lies on a Riemannian manifold. These indices are based on a Cramér von Mises like criterion that takes into account the geometry of the output support. We propose a Pick-Freeze like estimator of these indices based on an $U$--statistic. The asymptotic properties of these estimators are studied. Further, we provide and discuss some interesting numerical examples.

math.ST

Context tree selection for functional data

It has been repeatedly conjectured that the brain retrieves statistical regularities from stimuli. Here we present a new statistical approach allowing to address this conjecture. This approach is based on a new class of stochastic processes driven by chains with memory of variable length. It leads to a new experimental protocol in which sequences of auditory stimuli generated by a stochastic chain are presented to volunteers while electroencephalographic (EEG) data is recorded from their scalp. A new statistical model selection procedure for functional data is introduced and proved to be consistent. Applied to samples of EEG data collected using our experimental protocol it produces results supporting the conjecture that the brain effectively identifies the structure of the chain generating the sequence of stimuli.

q-bio.NC

Pattern recognition on random trees associated to protein functionality families

In this paper, we address the problem of identifying protein functionality using the information contained in its aminoacid sequence. We propose a method to define sequence similarity relationships that can be used as input for classification and clustering via well known metric based statistical methods. In our examples, we specifically address two problems of supervised and unsupervised learning in structural genomics via simple metric based techniques on the space of trees 1)Unsupervised detection of functionality families via K means clustering in the space of trees, 2)Classification of new proteins into known families via k nearest neighbour trees. We found evidence that the similarity measure induced by our approach concentrates information for discrimination. Classification has the same high performance than others VLMC approaches. Clustering is a harder task, though, but our approach for clustering is alignment free and automatic, and may lead to many interesting variations by choosing other clustering or classification procedures that are based on pre-computed similarity information, as the ones that performs clustering using flow simulation, see (Yona et al 2000, Enright et al, 2003).

stat.AP