arXiv · 1710.02616
Prediction analysis for microbiome sequencing data
Abstract
One primary goal of human microbiome studies is to predict host traits based on human microbiota. However, microbial community sequencing data present significant challenges to the development of statistical methods. In particular, the samples have different library sizes, the data contain many zeros and are often over-dispersed. To address these challenges, we introduce a new statistical framework, called predictive analysis in metagenomics via inverse regression (PAMIR). An inverse regression model is developed for over-dispersed microbiota counts given the trait, and then a prediction rule is constructed by taking advantage of the dimension-reduction structure in the model. An efficient Monte Carlo expectation-maximization algorithm is designed for carrying out maximum likelihood estimation. We demonstrate the advantages of PAMIR through simulations and a real data example.
Explore related subjects
Keep this discovery
Tao Wang, Can Yang, Hongyu Zhao. 2017-10-07. Prediction analysis for microbiome sequencing data. https://arxiv.org/abs/1710.02616
Cite the original work for its findings. Save a collection to share your selection of sources.