SearcharxivSearch

arXiv subjects

Terence P. Speed

Publications and source records attributed to Terence P. Speed.

6 recordsLinked to original sources

Building A Theoretical Foundation for Combining Negative Controls and Replicates

Studies using assays to quantify the expression of thousands of genes on tens to thousands of cell samples have been carried out for over 20 years. Such assays are based on microarrays, DNA sequencing or other molecular technologies. All such studies involve unwanted variation, often called batch effects, associated with the cell samples and the assay process. Removing this unwanted variation is essential before the measurements can be used to address the questions that motivated the studies. Combining the results of replicate assays with measurements on negative control genes to estimate the unwanted variation and remove it has proved to be effective at this task. The main goal of this paper is to present asymptotic theory that explains this effectiveness. The approach can be widened by using pseudo-replicate sets of pseudo-samples, for use with studies having no replicate assays. Theory covering this case is also presented. The established theory is supported by results of empirical investigations, including simulation studies and a real-data example.

math.ST

RLE Plots: Visualising Unwanted Variation in High Dimensional Data

Unwanted variation can be highly problematic and so its detection is often crucial. Relative log expression (RLE) plots are a powerful tool for visualising such variation in high dimensional data. We provide a detailed examination of these plots, with the aid of examples and simulation, explaining what they are and what they can reveal. RLE plots are particularly useful for assessing whether a procedure aimed at removing unwanted variation, i.e. a normalisation procedure, has been successful. These plots, while originally devised for gene expression data from microarrays, can also be used to reveal unwanted variation in many other kinds of high dimensional data, where such variation can be problematic.

stat.ME

Correcting gene expression data when neither the unwanted variation nor the factor of interest are observed

When dealing with large scale gene expression studies, observations are commonly contaminated by unwanted variation factors such as platforms or batches. Not taking this unwanted variation into account when analyzing the data can lead to spurious associations and to missing important signals. When the analysis is unsupervised, e.g., when the goal is to cluster the samples or to build a corrected version of the dataset - as opposed to the study of an observed factor of interest - taking unwanted variation into account can become a difficult task. The unwanted variation factors may be correlated with the unobserved factor of interest, so that correcting for the former can remove the latter if not done carefully. We show how negative control genes and replicate samples can be used to estimate unwanted variation in gene expression, and discuss how this information can be used to correct the expression data or build estimators for unsupervised problems. The proposed methods are then evaluated on three gene expression datasets. They generally manage to remove unwanted variation without losing the signal of interest and compare favorably to state of the art corrections.

stat.AP

Transcription factor binding site prediction with multivariate gene expression data

Multi-sample microarray experiments have become a standard experimental method for studying biological systems. A frequent goal in such studies is to unravel the regulatory relationships between genes. During the last few years, regression models have been proposed for the de novo discovery of cis-acting regulatory sequences using gene expression data. However, when applied to multi-sample experiments, existing regression based methods model each individual sample separately. To better capture the dynamic relationships in multi-sample microarray experiments, we propose a flexible method for the joint modeling of promoter sequence and multivariate expression data. In higher order eukaryotic genomes expression regulation usually involves combinatorial interaction between several transcription factors. Experiments have shown that spacing between transcription factor binding sites can significantly affect their strength in activating gene expression. We propose an adaptive model building procedure to capture such spacing dependent cis-acting regulatory modules. We apply our methods to the analysis of microarray time-course experiments in yeast and in Arabidopsis. These experiments exhibit very different dynamic temporal relationships. For both data sets, we have found all of the well-known cis-acting regulatory elements in the related context, as well as being able to predict novel elements.

stat.AP

Quality assessment for short oligonucleotide microarray data

Quality of microarray gene expression data has emerged as a new research topic. As in other areas, microarray quality is assessed by comparing suitable numerical summaries across microarrays, so that outliers and trends can be visualized, and poor quality arrays or variable quality sets of arrays can be identified. Since each single array comprises tens or hundreds of thousands of measurements, the challenge is to find numerical summaries which can be used to make accurate quality calls. To this end, several new quality measures are introduced based on probe level and probeset level information, all obtained as a by-product of the low-level analysis algorithms RMA/fitPLM for Affymetrix GeneChips. Quality landscapes spatially localize chip or hybridization problems. Numerical chip quality measures are derived from the distributions of Normalized Unscaled Standard Errors and of Relative Log Expressions. Quality of chip batches is assessed by Residual Scale Factors. These quality assessment measures are demonstrated on a variety of datasets (spike-in experiments, small lab experiments, multi-site studies). They are compared with Affymetrix's individual chip quality report.

stat.ME

A multivariate empirical Bayes statistic for replicated microarray time course data

In this paper we derive one- and two-sample multivariate empirical Bayes statistics (the $\mathit{MB}$-statistics) to rank genes in order of interest from longitudinal replicated developmental microarray time course experiments. We first use conjugate priors to develop our one-sample multivariate empirical Bayes framework for the null hypothesis that the expected temporal profile stays at 0. This leads to our one-sample $\mathit{MB}$-statistic and a one-sample $\widetilde{T}{}^2$-statistic, a variant of the one-sample Hotelling $T^2$-statistic. Both the $\mathit{MB}$-statistic and $\widetilde{T}^2$-statistic can be used to rank genes in the order of evidence of nonzero mean, incorporating the correlation structure across time points, moderation and replication. We also derive the corresponding $\mathit{MB}$-statistics and $\widetilde{T}^2$-statistics for the one-sample problem where the null hypothesis states that the expected temporal profile is constant, and for the two-sample problem where the null hypothesis is that two expected temporal profiles are the same.

math.ST