SearcharxivSearch

arXiv subjects

Tim J. P. Hubbard

Publications and source records attributed to Tim J. P. Hubbard.

3 recordsLinked to original sources

A Machine Learning Strategy to Identity Exonic Splice Enhancers in Human Protein-coding Sequence

Background: Exonic splice enhancers are sequences embedded within exons which promote and regulate the splicing of the transcript in which they are located. A class of exonic splice enhancers are the SR proteins, which are thought to mediate interactions between splicing factors bound to the 5' and 3' splice sites. Method and results: We present a novel strategy for analysing protein-coding sequence by first randomizing the codons used at each position within the coding sequence, then applying a motif-based machine learning algorithm to compare the true and randomized sequences. This strategy identified a collection of motifs which can successfully discriminate between real and randomized coding sequence, including -- but not restricted to -- several previously reported splice enhancer elements. As well as successfully distinguishing coding exons from randomized sequences, we show that our model is able to recognize non-coding exons. Conclusions: Our strategy succeeded in detecting signals in coding exons which seem to be orthogonal to the sequences' primary function of coding for proteins. We believe that many of the motifs detected here may represent binding sites for previously unrecognized proteins which influence RNA splicing. We hope that this development will lead to improved knowledge of exonic splice enhancers, and new developments in the field of computational gene prediction.

q-bio.GN

Relevance Vector Machines for classifying points and regions in biological sequences

The Relevance Vector Machine (RVM) is a recently developed machine learning framework capable of building simple models from large sets of candidate features. Here, we describe a protocol for using the RVM to explore very large numbers of candidate features, and a family of models which apply the power of the RVM to classifying and detecting interesting points and regions in biological sequence data. The models described here have been used successfully for predicting transcription start sites and other features in genome sequences.

q-bio.GN

What can we learn from noncoding regions of similarity between genomes?

Background: In addition to known protein-coding genes, large amount of apparently non-coding sequence are conserved between the human and mouse genomes. It seems reasonable to assume that these conserved regions are more likely to contain functional elements than less-conserved portions of the genome. Here we used a motif-oriented machine learning method to extract the strongest signal from a set of non-coding conserved sequences. Results: We successfully fitted models to reflect the non-coding sequences, and showed that the results were quite consistent for repeated training runs. Using the learned model to scan genomic sequence, we found that it often made predictions close to the start of annotated genes. We compared this method with other published promoter-prediction systems, and show that the set of promoters which are detected by this method seems to be substantially similar to that detected by existing methods. Conclusions: The results presented here indicate that the promoter signal is the strongest single motif-based signal in the non-coding functional fraction of the genome. They also lend support to the belief that there exists a substantial subset of promoter regions which share common features and are detectable by a variety of computational methods.

q-bio.GN