SearcharxivSearch

arXiv subjects

Ziliang Shen

Publications and source records attributed to Ziliang Shen.

4 recordsLinked to original sources

CoLaDAG: Compositional Latent Log-ratio DAG Analysis of the Gut Microbiome under Silver Nanoparticle Exposure

Directed network analysis of microbiome counts is complicated by compositional sampling, high dimensionality, and limited biological replication. We present CoLaDAG, a fixed-reference latent additive log-ratio (ALR) estimator for generating sparse directed conditional-dependence hypotheses from compositional counts. The method combines a multinomial observation model, a working linear Gaussian structural equation model, nonconvex DC-ADMM optimization, and post-estimation thresholding with greedy acyclic projection. Under simulations aligned with this observation model, CoLaDAG obtained the largest mean exact-direction Matthews correlation and the smallest mean false discovery rate among the evaluated implementations; performance deteriorated under continuous-data and dropout misspecification. In the 12-mouse silver-nanoparticle (AgNP) case study, the 58-node fitted graph was sensitive to block resampling and ALR reference choice: 60 of 284 primary edges attained a mouse-block selection frequency of at least 0.60. The reported orientations and dose-stratified slopes are exploratory, coordinate-specific hypotheses rather than identified causal or exposure effects. The leading stable relations prioritize anaerobic gut taxa for targeted abundance, metabolite, and perturbation studies, but do not establish cross-feeding or toxicological mechanisms.

stat.AP

High-Dimensional Differentially Private Quantile Regression: Distributed Estimation and Statistical Inference

With the development of big data and machine learning, privacy concerns have become increasingly critical, especially when handling heterogeneous datasets containing sensitive personal information. Differential privacy provides a rigorous framework for safeguarding individual privacy while enabling meaningful statistical analysis. In this paper, we propose a differentially private quantile regression method for high-dimensional data in a distributed setting. Quantile regression is a powerful and robust tool for modeling the relationships between the covariates and responses in the presence of outliers or heavy-tailed distributions. To address the computational challenges due to the non-smoothness of the quantile loss function, we introduce a Newton-type transformation that reformulates the quantile regression task into an ordinary least squares problem. Building on this, we develop a differentially private estimation algorithm with iterative updates, ensuring both near-optimal statistical accuracy and formal privacy guarantees. For inference, we further propose a differentially private debiased estimator, which enables valid confidence interval construction and hypothesis testing. Additionally, we propose a communication-efficient and differentially private bootstrap for simultaneous hypothesis testing in high-dimensional quantile regression, suitable for distributed settings with both small and abundant local data. Extensive simulations demonstrate the robustness and effectiveness of our methods in practical scenarios.

stat.ML

Sparsity learning via structured functional factor augmentation

As one of the most powerful tools for examining the association between functional covariates and a response, the functional regression model has been widely adopted in various interdisciplinary studies. Usually, a limited number of functional covariates are assumed in a functional linear regression model. Nevertheless, correlations may exist between functional covariates in high-dimensional functional linear regression models, which brings significant statistical challenges to statistical inference and functional variable selection. In this article, a novel functional factor augmentation structure (fFAS) is proposed for multivariate functional series, and a multivariate functional factor augmentation selection model (fFASM) is further proposed to deal with issues arising from variable selection of correlated functional covariates. Theoretical justifications for the proposed fFAS are provided, and statistical inference results of the proposed fFASM are established. Numerical investigations support the superb performance of the novel fFASM model in terms of estimation accuracy and selection consistency.

stat.ME

Distributed High-Dimensional Quantile Regression: Estimation Efficiency and Support Recovery

In this paper, we focus on distributed estimation and support recovery for high-dimensional linear quantile regression. Quantile regression is a popular alternative tool to the least squares regression for robustness against outliers and data heterogeneity. However, the non-smoothness of the check loss function poses big challenges to both computation and theory in the distributed setting. To tackle these problems, we transform the original quantile regression into the least-squares optimization. By applying a double-smoothing approach, we extend a previous Newton-type distributed approach without the restrictive independent assumption between the error term and covariates. An efficient algorithm is developed, which enjoys high computation and communication efficiency. Theoretically, the proposed distributed estimator achieves a near-oracle convergence rate and high support recovery accuracy after a constant number of iterations. Extensive experiments on synthetic examples and a real data application further demonstrate the effectiveness of the proposed method.

stat.ML