Searcharxiv⌕ Search

arXiv subjects

Michelle Pistner Nixon

Publications and source records attributed to Michelle Pistner Nixon.

3 recordsLinked to original sources

Scalable Bayesian Semiparametric Additive Regression Models For Microbiome Studies

Statistical analysis of microbiome data is challenging. Bayesian multinomial logistic-normal (MLN) models have gained popularity due to their ability to account for the count compositional nature of these data, but existing approaches are either computationally intractable or restricted to purely parametric or non-parametric methods, which limit their flexibility and scalability. In this work, we introduce \textit{MultiAddGPs}, a novel semi-parametric framework that integrates additive Gaussian Process (GP) regression within a Bayesian MLN model to disentangle linear and non-linear covariate effects, including non-stationary dynamics. Our approach builds on the computationally efficient Collapse-Uncollapse (CU) sampler and additive GP regression, introducing a novel back-sampling algorithm and marginal likelihood approximation for efficient inference and hyperparameter estimation. Our models are over 240,000 times faster than alternatives while simultaneously producing more accurate posterior estimates. Additionally, we incorporate non-stationary kernel functions designed to model treatment interventions and disease effects. We demonstrate our approach using simulated and real data studies and produce novel biological insights from a previously published human gut microbiome study. Our methods are publicly available as part of the \textit{fido} software package on CRAN \footnotemark.

stat.ME↗

Scale Reliant Inference

Many scientific fields, including human gut microbiome science, collect multivariate count data where the sum of the counts is unrelated to the scale of the underlying system being measured (e.g., total microbial load in a subject's colon). This disconnect complicates downstream analyses such as differential analysis in case-control studies. This article is motivated by a novel study of in vitro human gut microbiome models. Popular tools for analyzing these data led to dramatically elevated rates of both false positives and false negatives. To understand those failures, we provide a formal problem statement that frames these challenges of scale in terms of the classical theory of identifiability. We call this the problem of Scale Reliant Inference (SRI). We use this formulation to prove fundamental limits on SRI in terms of criteria such as consistency and type-I error control. We show that the failures of existing methods stem from a fundamental failure to properly quantify uncertainty in the system scale. We demonstrate that a particular type of Bayesian model called a Bayesian Partially Identified Model (PIMs) can correctly quantify uncertainty in SRI. We introduce Scale Simulation Random Variables (SSRVs) as a flexible and efficient approach to specifying and inferring Bayesian PIMs. In the context of both real and simulated data, we find SSRVs drastically decrease type-I and type-II error rates.

stat.ME↗

A Latent Class Modeling Approach for Generating Synthetic Data and Making Posterior Inferences from Differentially Private Counts

Several algorithms exist for creating differentially private counts from contingency tables, such as two-way or three-way marginal counts. The resulting noisy counts generally do not correspond to a coherent contingency table, so that some post-processing step is needed if one wants the released counts to correspond to a coherent contingency table. We present a latent class modeling approach for post-processing differentially private marginal counts that can be used (i) to create differentially private synthetic data from the set of marginal counts, and (ii) to enable posterior inferences about the confidential counts. We illustrate the approach using a subset of the 2016 American Community Survey Public Use Microdata Sets and the 2004 National Long Term Care Survey.

stat.ME↗