Searcharxiv⌕ Search

arXiv subjects

Dongyue Xie

Publications and source records attributed to Dongyue Xie.

4 recordsLinked to original sources

A Flexible Empirical Bayes Approach to Generalized Linear Models, with Applications to Sparse Logistic Regression

We introduce a flexible empirical Bayes approach for fitting Bayesian generalized linear models. Specifically, we adopt a novel mean-field variational inference (VI) method and the prior is estimated within the VI algorithm, making the method tuning-free. Unlike traditional VI methods that optimize the posterior density function, our approach directly optimizes the posterior mean and prior parameters. This formulation reduces the number of parameters to optimize and enables the use of scalable algorithms such as L-BFGS and stochastic gradient descent. Furthermore, our method automatically determines the optimal posterior based on the prior and likelihood, distinguishing it from existing VI methods that often assume a Gaussian variational. Our approach represents a unified framework applicable to a wide range of exponential family distributions, removing the need to develop unique VI methods for each combination of likelihood and prior distributions. We apply the framework to solve sparse logistic regression and demonstrate the superior predictive performance of our method in extensive numerical studies, by comparing it to prevalent sparse logistic regression approaches.

stat.ML↗

Statistical Inference for Cell Type Deconvolution

Integrating heterogeneous datasets across different measurement platforms is a fundamental challenge in many scientific applications. A common example arises in deconvolution problems, such as cell type deconvolution, where one aims to estimate the composition of latent subpopulations using reference data from a different source. However, this task is complicated by systematic platform-specific scaling effects, measurement noise, and differences between data sources. For the problem of cell type deconvolution, existing methods often neglect the correlation and uncertainty in cell type proportion estimates, possibly leading to an additional concern of false positives in downstream comparisons across multiple individuals. We introduce MEAD, a statistical framework that provides both accurate estimation and valid statistical inference on the estimates. One of our key contributions is the identifiability result, which establishes the conditions under which cell type compositions are identifiable under arbitrary gene-specific scaling differences across platforms. MEAD also supports the comparison of cell type proportions across individuals after deconvolution, accounting for gene-gene correlations and biological variability. Through simulations and real-data analysis, MEAD demonstrates superior reliability for inferring cell type compositions in complex biological systems.

stat.ME↗

Local False Sign Rate and the Role of Prior Covariance Rank in Multivariate Empirical Bayes Multiple Testing

This paper investigates the relationship between the rank of the prior covariance matrix and the local false sign rate (lfsr) in multivariate empirical Bayes multiple testing, specifically within the context of normal mean models. We demonstrate that using low-rank covariance matrices for the prior results in inflated false sign rates, a consequence of rank deficiency. To address this, we propose an adjustment that mitigates this inflation by employing full-rank covariance matrices. Through simulations, we validate the effectiveness of this adjustment in controlling false sign rates, thereby improving the robustness of empirical Bayes methods in high-dimensional settings. Our results show that the rank of the prior covariance matrix directly influences the accuracy of sign estimation and the performance of the lfsr, with significant implications for large-scale hypothesis testing in statistics and genomics.

stat.ME↗

A flexible model for correlated count data, with application to multi-condition differential expression analyses of single-cell RNA sequencing data

Detecting differences in gene expression is an important part of single-cell RNA sequencing experiments, and many statistical methods have been developed for this aim. Most differential expression analyses focus on comparing expression between two groups (e.g., treatment vs. control). But there is increasing interest in multi-condition differential expression analyses in which expression is measured in many conditions, and the aim is to accurately detect and estimate expression differences in all conditions. We show that directly modeling single-cell RNA-seq counts in all conditions simultaneously, while also inferring how expression differences are shared across conditions, leads to greatly improved performance for detecting and estimating expression differences compared to existing methods. We illustrate the potential of this new approach by analyzing data from a single-cell experiment studying the effects of cytokine stimulation on gene expression. We call our new method "Poisson multivariate adaptive shrinkage", and it is implemented in an R package available online at https://github.com/stephenslab/poisson.mash.alpha.

stat.ME↗