Searcharxiv⌕ Search

arXiv subjects

Jerome Saracco

Publications and source records attributed to Jerome Saracco.

3 recordsLinked to original sources

Study of inter-individual variability of three-dimensional data table: detection of unstable variables and samples

We propose two methodologies in order to better understand the inter-individual variability of resting-state functional Magnetic Resonance Imaging (fMRI) brain data. The aim of the study was to quantify whether the average dendrogram is representative of the initial population and to identify its possible sources of instability. The average dendrogram is based on the Pearson correlation between resting-state networks. The first method identifies networks that can lead to unstable partitions of the average dendrogram. The second method identified homogeneous sub-samples of participants for whom their associated average dendrograms were more stable than that of the whole sample. The two suggested methods have shown significant quantifiable behavioral data results with regards to detecting an unstable network or presence of subpopulations when the noise level does not conceal the structure of the data. These two methods have been successfully applied to establish a cerebral atlas for late adulthood. The first method made it clear that there was no unstable network among the atlas networks. The second method highlighted the presence of two distinct sub-populations with different age-related brain organizations.

physics.med-ph↗

Combining clustering of variables and feature selection using random forests

Standard approaches to tackle high-dimensional supervised classification problem often include variable selection and dimension reduction procedures. The novel methodology proposed in this paper combines clustering of variables and feature selection. More precisely, hierarchical clustering of variables procedure allows to build groups of correlated variables in order to reduce the redundancy of information and summarizes each group by a synthetic numerical variable. Originality is that the groups of variables (and the number of groups) are unknown a priori. Moreover the clustering approach used can deal with both numerical and categorical variables (i.e. mixed dataset). Among all the possible partitions resulting from dendrogram cuts, the most relevant synthetic variables (i.e. groups of variables) are selected with a variable selection procedure using random forests. Numerical performances of the proposed approach are compared with direct applications of random forests and variable selection using random forests on the original p variables. Improvements obtained with the proposed methodology are illustrated on two simulated mixed datasets (cases n>p and n<p, where n is the sample size) and on a real proteomic dataset. Via the selection of groups of variables (based on the synthetic variables), interpretability of the results becomes easier.

math.ST↗

On the asymptotic behavior of the Nadaraya-Watson estimator associated with the recursive SIR method

We investigate the asymptotic behavior of the Nadaraya-Watson estimator for the estimation of the regression function in a semiparametric regression model. On the one hand, we make use of the recursive version of the sliced inverse regression method for the estimation of the unknown parameter of the model. On the other hand, we implement a recursive Nadaraya-Watson procedure for the estimation of the regression function which takes into account the previous estimation of the parameter of the semiparametric regression model. We establish the almost sure convergence as well as the asymptotic normality for our Nadaraya-Watson estimator. We also illustrate our semiparametric estimation procedure on simulated data.

math.ST↗