SearcharxivSearch

arXiv subjects

Luke Jostins

Publications and source records attributed to Luke Jostins.

3 recordsLinked to original sources

Latent variable model selection for Gaussian conditional random fields

We consider the problem of learning a conditional Gaussian graphical model in the presence of latent variables. Building on recent advances in this field, we suggest a method that decomposes the parameters of a conditional Markov random field into the sum of a sparse and a low-rank matrix. We derive convergence bounds for this estimator and show that it is well-behaved in the high-dimensional regime as well as "sparsistent" (i.e. capable of recovering the graph structure). We then show how proximal gradient algorithms and semi-definite programming techniques can be employed to fit the model to thousands of variables. Through extensive simulations, we illustrate the conditions required for identifiability and show that there is a wide range of situations in which this model performs significantly better than its counterparts, for example, by accommodating more latent variables. Finally, the suggested method is applied to two datasets comprising individual level data on genetic variants and metabolites levels. We show our results replicate better than alternative approaches and show enriched biological signal.

stat.ME

YFitter: Maximum likelihood assignment of Y chromosome haplogroups from low-coverage sequence data

Low-coverage short-read resequencing experiments have the potential to expand our understanding of Y chromosome haplogroups. However, the uncertainty associated with these experiments mean that haplogroups must be assigned probabilistically to avoid false inferences. We propose an efficient dynamic programming algorithm that can assign haplogroups by maximum likelihood, and represent the uncertainty in assignment. We apply this to both genotype and low-coverage sequencing data, and show that it can assign haplogroups accurately and with high resolution. The method is implemented as the program YFitter, which can be downloaded from http://sourceforge.net/projects/yfitter/

q-bio.PE

Inferring genotyping error rates from genotyped trios

Genotyping errors are known to influence the power of both family-based and case-control studies in the genetics of complex disease. Estimating genotyping error rate in a given dataset can be complex, but when family information is available error rates can be inferred from the patterns of Mendelian inheritance between parents and offspring. I introduce a novel likelihood-based method for calculating error rates from family data, given known allele frequencies. I apply this to an example dataset, demonstrating a low genotyping error rate in genotyping data from a personal genomics company.

q-bio.QM