SearcharxivSearch

arXiv subjects

Shuangjie Zhang

Publications and source records attributed to Shuangjie Zhang.

3 recordsLinked to original sources

Structure Learning for Directed Trees with Zero-Inflated Compositional Nodes

Compositional data, which are vectors of proportions constrained to the probability simplex, arise frequently in modern scientific applications, including microbiome relative abundances across body sites and cell-type mixture weights derived from single-cell genomics. While regression methods for compositional data are well developed, no existing graphical model framework addresses the problem of learning conditional dependence structures among multiple compositional vectors. This paper introduces a novel framework for directed tree structure learning over compositional nodes. We employ the Kullback-Leibler divergence as the scoring function and model the conditional expectation of each child composition as a mixture of a baseline composition and a parent-driven component parameterized by a column-stochastic transition matrix. This formulation respects the simplex geometry, handles zero-inflated compositions gracefully, and, combined with a non-degeneracy condition on the transition matrix, ensures identifiability of edge directions from observational data. We prove consistency of structure recovery and derive finite-sample guarantees that characterize the required sample size in terms of the signal gap, node dimension, and penalty level. The efficacy of our approach is demonstrated through simulations and applications to multi-site microbiome data and single-cell data, yielding interpretable directed structures that align with known biological mechanisms.

stat.ME

Bayesian Covariate-Varying Interaction Analysis for Multivariate Count Data: Application to Microbiome Studies

Understanding covariate-varying interdependencies among features is of great interest in various applications. Motivated by microbiome studies where microbial abundances and interactions vary with environmental factors, we develop a Bayesian covariate-varying factor model. This model flexibly estimates heteroscedasticity in the covariance matrix as a function of covariates. Specifically, our approach employs covariance regression through linear regression on a lower-dimensional factor loading matrix. This formulation, combined with joint sparsity induced by the Dirichlet--Horseshoe prior for the factor loadings, provides robust estimation of covariate-varying covariance in high-dimensional settings. The model simultaneously incorporates a regression structure for the mean abundance and jointly addresses the covariate-varying mean and covariance structure. Furthermore, the model tackles key statistical challenges such as discreteness, over-dispersion, compositionality, and high dimensionality, common in microbiome data analysis, using a flexible nonparametric Bayesian framework. We thoroughly investigate the properties of the model and conduct extensive simulation studies to examine its performance. Real microbiome data examples are provided for illustration.

stat.ME

Density Discontinuity Regression

Many policies hinge on a continuous variable exceeding a threshold, prompting strategic behavior by agents to stay on the favorable side. This creates density discontinuities at cutoffs, evident in contexts like taxable income, corporate regulations, and academic grading. Existing methods detect these discontinuities, but systematic approaches to examine how they vary with observable characteristics are lacking. We propose a novel, interpretable Bayesian framework that jointly estimates both the log-density ratio at the cutoff and the local shape of the density, as functions of covariates, within a data-driven window. This formulation yields regression-style estimates of covariate effects on the discontinuity. An adaptive window selection balances bias and variance. Our approach improves upon common methods that target only the log-density ratio around the threshold while ignoring the local density shape. We constrain the density jump to be non-negative, reflecting that agents would not aim to be on the losing side of the threshold. Applied to corporate shareholder voting data, our method identifies substantial variation in strategic behavior, notably stronger discontinuities for proposals facing negative recommendations from Institutional Shareholder Services, larger firms, and firms with lower analyst coverage. Overall, our method provides an interpretable framework to quantify heterogeneous agent responses to threshold-based policies.

stat.ME