SearcharxivSearch

arXiv subjects

Yves Rozenholc

Publications and source records attributed to Yves Rozenholc.

14 recordsLinked to original sources

Differential analysis in Transcriptomic: The strength of randomly picking 'reference' genes

Transcriptomic analysis are characterized by being not directly quantitative and only providing relative measurements of expression levels up to an unknown individual scaling factor. This difficulty is enhanced for differential expression analysis. Several methods have been proposed to circumvent this lack of knowledge by estimating the unknown individual scaling factors however, even the most used one, are suffering from being built on hardly justifiable biological hypotheses or from having weak statistical background. Only two methods withstand this analysis: one based on largest connected graph component hardly usable for large amount of expressions like in NGS, the second based on $\log$-linear fits which unfortunately require a first step which uses one of the methods described before. We introduce a new procedure for differential analysis in the context of transcriptomic data. It is the result of pooling together several differential analyses each based on randomly picked genes used as reference genes. It provides a differential analysis free from the estimation of the individual scaling factors or any other knowledge. Theoretical properties are investigated both in term of FWER and power. Moreover in the context of Poisson or negative binomial modelization of the transcriptomic expressions, we derived a test with non asymptotic control of its bounds. We complete our study by some empirical simulations and apply our procedure to a real data set of hepatic miRNA expressions from a mouse model of non-alcoholic steatohepatitis (NASH), the CDAHFD model. This study on real data provides new hits with good biological explanations.

stat.ME

Hierarchical segmentation using equivalence test (HiSET): Application to DCE image sequences

Dynamical contrast enhanced (DCE) imaging allows non invasive access to tissue micro-vascularization. It appears as a promising tool to build imaging biomark-ers for diagnostic, prognosis or anti-angiogenesis treatment monitoring of cancer. However, quantitative analysis of DCE image sequences suffers from low signal to noise ratio (SNR). SNR may be improved by averaging functional information in a large region of interest when it is functionally homogeneous. We propose a novel method for automatic segmentation of DCE image sequences into functionally homogeneous regions, called DCE-HiSET. Using an observation model which depends on one parameter a and is justified a posteri-ori, DCE-HiSET is a hierarchical clustering algorithm. It uses the p-value of a multiple equivalence test as dissimilarity measure and consists of two steps. The first exploits the spatial neighborhood structure to reduce complexity and takes advantage of the regularity of anatomical features, while the second recovers (spatially) disconnected homogeneous structures at a larger (global) scale. Given a minimal expected homogeneity discrepancy for the multiple equivalence test, both steps stop automatically by controlling the Type I error. This provides an adaptive choice for the number of clusters. Assuming that the DCE image sequence is functionally piecewise constant with signals on each piece sufficiently separated, we prove that DCE-HiSET will retrieve the exact partition with high probability as soon as the number of images in the sequence is large enough. The minimal expected homogeneity discrepancy appears as the tuning parameter controlling the size of the segmentation. DCE-HiSET has been implemented in C++ for 2D and 3D image sequences with competitive speed. Keywords : DCE imaging, automatic clustering, hierarchical segmentation, equivalence test

eess.IV

Laplace deconvolution on the basis of time domain data and its application to Dynamic Contrast Enhanced imaging

In the present paper we consider the problem of Laplace deconvolution with noisy discrete non-equally spaced observations on a finite time interval. We propose a new method for Laplace deconvolution which is based on expansions of the convolution kernel, the unknown function and the observed signal over Laguerre functions basis (which acts as a surrogate eigenfunction basis of the Laplace convolution operator) using regression setting. The expansion results in a small system of linear equations with the matrix of the system being triangular and Toeplitz. Due to this triangular structure, there is a common number $m$ of terms in the function expansions to control, which is realized via complexity penalty. The advantage of this methodology is that it leads to very fast computations, produces no boundary effects due to extension at zero and cut-off at $T$ and provides an estimator with the risk within a logarithmic factor of the oracle risk. We emphasize that, in the present paper, we consider the true observational model with possibly nonequispaced observations which are available on a finite interval of length $T$ which appears in many different contexts, and account for the bias associated with this model (which is not present when $T\rightarrow\infty$). The study is motivated by perfusion imaging using a short injection of contrast agent, a procedure which is applied for medical assessment of micro-circulation within tissues such as cancerous tumors. Presence of a tuning parameter $a$ allows to choose the most advantageous time units, so that both the kernel and the unknown right hand side of the equation are well represented for the deconvolution. The methodology is illustrated by an extensive simulation study and a real data example which confirms that the proposed technique is fast, efficient, accurate, usable from a practical point of view and very competitive.

stat.ME

An efficient algorithm for T-estimation

We introduce an efficient and exact algorithm, together with a faster but approximate version, which implements with a sub-quadratic complexity the hold-out derived from T-estimation. We study empirically the performance of this hold-out in the context of density estimation considering well-known competitors (hold-out derived from least-squares or Kullback-Leibler divergence, model selection procedures, etc.) and classical problems including histogram or bandwidth selection. Our algorithms are integrated in a companion R-package called {\it Density.T.HoldOut} available on the CRAN: {\url{http://cran.r-project.org/web/packages/Density.T.HoldOut/index.html}}.

stat.CO

Fast estimation of posterior probabilities in change-point models through a constrained hidden Markov model

The detection of change-points in heterogeneous sequences is a statistical challenge with applications across a wide variety of fields. In bioinformatics, a vast amount of methodology exists to identify an ideal set of change-points for detecting Copy Number Variation (CNV). While considerable efficient algorithms are currently available for finding the best segmentation of the data in CNV, relatively few approaches consider the important problem of assessing the uncertainty of the change-point location. Asymptotic and stochastic approaches exist but often require additional model assumptions to speed up the computations, while exact methods have quadratic complexity which usually are intractable for large datasets of tens of thousands points or more. In this paper, we suggest an exact method for obtaining the posterior distribution of change-points with linear complexity, based on a constrained hidden Markov model. The methods are implemented in the R package postCP, which uses the results of a given change-point detection algorithm to estimate the probability that each observation is a change-point. We present the results of the package on a publicly available CNV data set (n=120). Due to its frequentist framework, postCP obtains less conservative confidence intervals than previously published Bayesian methods, but with linear complexity instead of quadratic. Simulations showed that postCP provided comparable loss to a Bayesian MCMC method when estimating posterior means, specifically when assessing larger-scale changes, while being more computationally efficient. On another high-resolution CNV data set (n=14,241), the implementation processed information in less than one second on a mid-range laptop computer.

stat.AP

Laplace deconvolution with noisy observations

In the present paper we consider Laplace deconvolution for discrete noisy data observed on the interval whose length may increase with a sample size. Although this problem arises in a variety of applications, to the best of our knowledge, it has been given very little attention by the statistical community. Our objective is to fill this gap and provide statistical treatment of Laplace deconvolution problem with noisy discrete data. The main contribution of the paper is explicit construction of an asymptotically rate-optimal (in the minimax sense) Laplace deconvolution estimator which is adaptive to the regularity of the unknown function. We show that the original Laplace deconvolution problem can be reduced to nonparametric estimation of a regression function and its derivatives on the interval of growing length T_n. Whereas the forms of the estimators remain standard, the choices of the parameters and the minimax convergence rates, which are expressed in terms of T_n^2/n in this case, are affected by the asymptotic growth of the length of the interval. We derive an adaptive kernel estimator of the function of interest, and establish its asymptotic minimaxity over a range of Sobolev classes. We illustrate the theory by examples of construction of explicit expressions of Laplace deconvolution estimators. A simulation study shows that, in addition to providing asymptotic optimality as the number of observations turns to infinity, the proposed estimator demonstrates good performance in finite sample examples.

math.ST

Generalization of the normal-exponential model: exploration of a more accurate parametrisation for the signal distribution on Illumina BeadArrays

Motivation: Illumina BeadArray technology includes negative control features that allow a precise estimation of the background noise. As an alternative to the background subtraction proposed in BeadStudio which leads to an important loss of information by generating negative values, a background correction method modeling the observed intensities as the sum of the exponentially distributed signal and normally distributed noise has been developed. Nevertheless, Wang and Ye (2011) display a kernel-based estimator of the signal distribution on Illumina BeadArrays and suggest that a gamma distribution would represent a better modeling of the signal density. Hence, the normal-exponential modeling may not be appropriate for Illumina data and background corrections derived from this model may lead to wrong estimation. Results: We propose a more flexible modeling based on a gamma distributed signal and a normal distributed background noise and develop the associated background correction. Our model proves to be markedly more accurate to model Illumina BeadArrays: on the one hand, this model offers a more correct fit of the observed intensities. On the other hand, the comparison of the operating characteristics of several background correction procedures on spike-in and on normal-gamma simulated data shows high similarities, reinforcing the validation of the normal-gamma modeling. The performance of the background corrections based on the normal-gamma and normal-exponential models are compared on two dilution data sets. Surprisingly, we observe that the implementation of a more accurate parametrisation in the model-based background correction does not increase the sensitivity. These results may be explained by the operating characteristics of the estimators: the normal-gamma background correction offers an improvement in terms of bias, but at the cost of a loss in precision.

stat.AP

Laplace deconvolution and its application to Dynamic Contrast Enhanced imaging

In the present paper we consider the problem of Laplace deconvolution with noisy discrete observations. The study is motivated by Dynamic Contrast Enhanced imaging using a bolus of contrast agent, a procedure which allows considerable improvement in {evaluating} the quality of a vascular network and its permeability and is widely used in medical assessment of brain flows or cancerous tumors. Although the study is motivated by medical imaging application, we obtain a solution of a general problem of Laplace deconvolution based on noisy data which appears in many different contexts. We propose a new method for Laplace deconvolution which is based on expansions of the convolution kernel, the unknown function and the observed signal over Laguerre functions basis. The expansion results in a small system of linear equations with the matrix of the system being triangular and Toeplitz. The number $m$ of the terms in the expansion of the estimator is controlled via complexity penalty. The advantage of this methodology is that it leads to very fast computations, does not require exact knowledge of the kernel and produces no boundary effects due to extension at zero and cut-off at $T$. The technique leads to an estimator with the risk within a logarithmic factor of $m$ of the oracle risk under no assumptions on the model and within a constant factor of the oracle risk under mild assumptions. The methodology is illustrated by a finite sample simulation study which includes an example of the kernel obtained in the real life DCE experiments. Simulations confirm that the proposed technique is fast, efficient, accurate, usable from a practical point of view and competitive.

math.ST

Craniofacial reconstruction as a prediction problem using a Latent Root Regression model

In this paper, we present a computer-assisted method for facial reconstruction. This method provides an estimation of the facial shape associated with unidentified skeletal remains. Current computer-assisted methods using a statistical framework rely on a common set of extracted points located on the bone and soft-tissue surfaces. Most of the facial reconstruction methods then consist of predicting the position of the soft-tissue surface points, when the positions of the bone surface points are known. We propose to use Latent Root Regression for prediction. The results obtained are then compared to those given by Principal Components Analysis linear models. In conjunction, we have evaluated the influence of the number of skull landmarks used. Anatomical skull landmarks are completed iteratively by points located upon geodesics which link these anatomical landmarks, thus enabling us to artificially increase the number of skull points. Facial points are obtained using a mesh-matching algorithm between a common reference mesh and individual soft-tissue surface meshes. The proposed method is validated in term of accuracy, based on a leave-one-out cross-validation test applied to a homogeneous database. Accuracy measures are obtained by computing the distance between the original face surface and its reconstruction. Finally, these results are discussed referring to current computer-assisted reconstruction facial techniques.

cs.LG

Pointwise adaptive estimation for robust and quantile regression

A nonparametric procedure for robust regression estimation and for quantile regression is proposed which is completely data-driven and adapts locally to the regularity of the regression function. This is achieved by considering in each point M-estimators over different local neighbourhoods and by a local model selection procedure based on sequential testing. Non-asymptotic risk bounds are obtained, which yield rate-optimality for large sample asymptotics under weak conditions. Simulations for different univariate median regression models show good finite sample properties, also in comparison to traditional methods. The approach is extended to image denoising and applied to CT scans in cancer research.

math.ST

Nonparametric estimation for a stochastic volatility model

Consider discrete time observations (X_{\ellδ})_{1\leq \ell \leq n+1}$ of the process $X$ satisfying $dX_t= \sqrt{V_t} dB_t$, with $V_t$ a one-dimensional positive diffusion process independent of the Brownian motion $B$. For both the drift and the diffusion coefficient of the unobserved diffusion $V$, we propose nonparametric least square estimators, and provide bounds for theirrisk. Estimators are chosen among a collection of functions belonging to a finite dimensional space whose dimension is selected by a data driven procedure. Implementation on simulated data illustrates how the method works.

stat.ME

Penalized nonparametric mean square estimation of the coefficients of diffusion processes

We consider a one-dimensional diffusion process $(X_t)$ which is observed at $n+1$ discrete times with regular sampling interval $Δ$. Assuming that $(X_t)$ is strictly stationary, we propose nonparametric estimators of the drift and diffusion coefficients obtained by a penalized least squares approach. Our estimators belong to a finite-dimensional function space whose dimension is selected by a data-driven method. We provide non-asymptotic risk bounds for the estimators. When the sampling interval tends to zero while the number of observations and the length of the observation time interval tend to infinity, we show that our estimators reach the minimax optimal rates of convergence. Numerical results based on exact simulations of diffusion processes are given for several examples of models and illustrate the qualities of our estimation algorithms.

math.ST

Penalized contrast estimator for adaptive density deconvolution

The authors consider the problem of estimating the density $g$ of independent and identically distributed variables $X\_i$, from a sample $Z\_1, ..., Z\_n$ where $Z\_i=X\_i+σε\_i$, $i=1, ..., n$, $ε$ is a noise independent of $X$, with $σε$ having known distribution. They present a model selection procedure allowing to construct an adaptive estimator of $g$ and to find non-asymptotic bounds for its $\mathbb{L}\_2(\mathbb{R})$-risk. The estimator achieves the minimax rate of convergence, in most cases where lowers bounds are available. A simulation study gives an illustration of the good practical performances of the method.

math.ST

Finite sample penalization in adaptive density deconvolution

We consider the problem of estimating the density $g$ of identically distributed variables $X\_i$, from a sample $Z\_1, ..., Z\_n$ where $Z\_i=X\_i+σε\_i$, $i=1, ..., n$ and $σε\_i$ is a noise independent of $X\_i$ with known density $ σ^{-1}f\_ε(./σ)$. We generalize adaptive estimators, constructed by a model selection procedure, described in Comte et al. (2005). We study numerically their properties in various contexts and we test their robustness. Comparisons are made with respect to deconvolution kernel estimators, misspecification of errors, dependency,... It appears that our estimation algorithm, based on a fast procedure, performs very well in all contexts.

math.ST