SearcharxivSearch

arXiv subjects

Ruibin Xi

Publications and source records attributed to Ruibin Xi.

9 recordsLinked to original sources

Distribution-free screening of spatially variable genes in spatial transcriptomics

Spatial transcriptomics (ST) technologies enable transcriptome-wide gene expression profiling while preserving spatial resolution, offering unprecedented opportunities to uncover complex spatial structures. Due to the ultra-high dimensionality of ST data, identifying spatially variable genes (SVGs) associated with unknown spatial clusters has become a central task in ST data analysis. Here, we develop a distribution-free SVG screening method based on a novel quasi-likelihood ratio statistic, the MM-test, combined with a knockoff procedure to control the false discovery rate (FDR). MM-test leverages auxiliary information, such as spatial distances, about the unknown spatial domains for SVG screening. Notably, in addition to two-dimensional ST datasets, MM-test is well-suited for increasingly common three-dimensional (3D), multi-slice ST datasets. Extensive benchmarking using simulations and 34 real ST datasets demonstrates that MM-test consistently outperforms existing SVG detection methods. In a 3D mouse brain dataset, MM-test accurately delineates fine-scale structures that are challenging for other methods, such as the 3D architecture of the pyramidal layer of the hippocampal cornu ammonis and the dentate gyrus. Theoretical guarantees-including selection consistency, FDR control, and an error bound for post-selection clustering-are also established.

stat.AP

Generic properties of vector fields identical on a compact set and codimension one partially hyperbolic dynamics

Let $\mathscr{X}^r(M)$ be the set of $C^r$ vector fields on a boundaryless compact Riemannian manifold $M$. Given a vector field $X_0\in\mathscr{X}^r(M)$ and a compact invariant set $\Gamma$ of $X_0$, we consider the closed subset $\mathscr{X}^r(M,\Gamma)$ of $\mathscr{X}^r(M)$, consisting of all $C^r$ vector fields which coincide with $X_0$ on $\Gamma$. Study of such a set naturally arises when one needs to perturb a system while keeping part of the dynamics untouched. A vector field $X\in\mathscr{X}^r(M,\Gamma)$ is called $\Gamma$-avoiding Kupka-Smale, if the dynamics away from $\Gamma$ is Kupka-Smale. We show that a generic vector field in $\mathscr{X}^r(M,\Gamma)$ is $\Gamma$-avoiding Kupka-Smale. In the $C^1$ topology, we obtain more generic properties for $\mathscr{X}^1(M,\Gamma)$. With these results, we further study codimension one partially hyperbolic dynamics for generic vector fields in $\mathscr{X}^1(M,\Gamma)$, giving a dichotomy of hyperbolicity and Newhouse phenomenon. As an application, we obtain that $C^1$ generically in $\mathscr{X}^1(M)$, a non-trivial Lyapunov stable chain recurrence class of a singularity which admits a codimension 2 partially hyperbolic splitting with respect to the tangent flow is a homoclinic class.

math.DS

Feature screening for clustering analysis

In this paper, we consider feature screening for ultrahigh dimensional clustering analyses. Based on the observation that the marginal distribution of any given feature is a mixture of its conditional distributions in different clusters, we propose to screen clustering features by independently evaluating the homogeneity of each feature's mixture distribution. Important cluster-relevant features have heterogeneous components in their mixture distributions and unimportant features have homogeneous components. The well-known EM-test statistic is used to evaluate the homogeneity. Under general parametric settings, we establish the tail probability bounds of the EM-test statistic for the homogeneous and heterogeneous features, and further show that the proposed screening procedure can achieve the sure independent screening and even the consistency in selection properties. Limiting distribution of the EM-test statistic is also obtained for general parametric distributions. The proposed method is computationally efficient, can accurately screen for important cluster-relevant features and help to significantly improve clustering, as demonstrated in our extensive simulation and real data analyses.

stat.ME

Network Analysis of Count Data from Mixed Populations

In applications such as gene regulatory network analysis based on single-cell RNA sequencing data, samples often come from a mixture of different populations and each population has its own unique network. Available graphical models often assume that all samples are from the same population and share the same network. One has to first cluster the samples and use available methods to infer the network for every cluster separately. However, this two-step procedure ignores uncertainty in the clustering step and thus could lead to inaccurate network estimation. Motivated by these applications, we consider the mixture Poisson log-normal model for network inference of count data from mixed populations. The latent precision matrices of the mixture model correspond to the networks of different populations and can be jointly estimated by maximizing the lasso-penalized log-likelihood. Under rather mild conditions, we show that the mixture Poisson log-normal model is identifiable and has the positive definite Fisher information matrix. Consistency of the maximum lasso-penalized log-likelihood estimator is also established. To avoid the intractable optimization of the log-likelihood, we develop an algorithm called VMPLN based on the variational inference method. Comprehensive simulation and real single-cell RNA sequencing data analyses demonstrate the superior performance of VMPLN.

stat.ME

Single-cell gene regulatory network analysis for mixed cell populations with applications to COVID-19 single cell data

Gene regulatory network (GRN) refers to the complex network formed by regulatory interactions between genes in living cells. In this paper, we consider inferring GRNs in single cells based on single cell RNA sequencing (scRNA-seq) data. In scRNA-seq, single cells are often profiled from mixed populations and their cell identities are unknown. A common practice for single cell GRN analysis is to first cluster the cells and infer GRNs for every cluster separately. However, this two-step procedure ignores uncertainty in the clustering step and thus could lead to inaccurate estimation of the networks. To address this problem, we propose to model scRNA-seq by the mixture multivariate Poisson log-normal (MPLN) distribution. The precision matrices of the MPLN are the GRNs of different cell types and can be jointly estimated by maximizing MPLN's lasso-penalized log-likelihood. We show that the MPLN model is identifiable and the resulting penalized log-likelihood estimator is consistent. To avoid the intractable optimization of the MPLN's log-likelihood, we develop an algorithm called VMPLN based on the variational inference method. Comprehensive simulation and real scRNA-seq data analyses reveal that VMPLN performs better than the state-of-the-art single cell GRN methods.

q-bio.MN

Gene regulatory network in single cells based on the Poisson log-normal model

Gene regulatory network inference is crucial for understanding the complex molecular interactions in various genetic and environmental conditions. The rapid development of single-cell RNA sequencing (scRNA-seq) technologies unprecedentedly enables gene regulatory networks inference at the single cell resolution. However, traditional graphical models for continuous data, such as Gaussian graphical models, are inappropriate for network inference of scRNA-seq's count data. Here, we model the scRNA-seq data using the multivariate Poisson log-normal (PLN) distribution and represent the precision matrix of the latent normal distribution as the regulatory network. We propose to first estimate the latent covariance matrix using a moment estimator and then estimate the precision matrix by minimizing the lasso-penalized D-trace loss function. We establish the convergence rate of the covariance matrix estimator and further establish the convergence rates and the sign consistency of the proposed PLNet estimator of the precision matrix in the high dimensional setting. The performance of PLNet is evaluated and compared with available methods using simulation and gene regulatory network analysis of scRNA-seq data.

stat.ME

A robust statistical method for Genome-wide association analysis of human copy number variation

Conducting genome-wide association studies (GWAS) in copy number variation (CNV) level is a field where few people involves and little statistical progresses have been achieved, traditional methods suffer from many problems such as batch effects, heterogeneity across genome, leading to low power or high false discovery rate. We develop a new robust method to find disease-risking regions related to CNV's disproportionately distributed between case and control samples, even if there are batch effects between them, our test formula is robust to such effects. We propose a new empirical Bayes rule to deal with overfitting when estimating parameters during testing, this rule can be extended to the field of model selection, it can be more efficient compared with traditional methods when there are too much potential models to be specified. We also give solid theoretical guarantees for our proposed method, and demonstrate the effectiveness by simulation and realdata analysis.

stat.ME

Community Detection by $L_0$-penalized Graph Laplacian

Community detection in network analysis aims at partitioning nodes in a network into $K$ disjoint communities. Most currently available algorithms assume that $K$ is known, but choosing a correct $K$ is generally very difficult for real networks. In addition, many real networks contain outlier nodes not belonging to any community, but currently very few algorithm can handle networks with outliers. In this paper, we propose a novel model free tightness criterion and an efficient algorithm to maximize this criterion for community detection. This tightness criterion is closely related with the graph Laplacian with $L_0$ penalty. Unlike most community detection methods, our method does not require a known $K$ and can properly detect communities in networks with outliers. Both theoretical and numerical properties of the method are analyzed. The theoretical result guarantees that, under the degree corrected stochastic block model, even for networks with outliers, the maximizer of the tightness criterion can extract communities with small misclassification rates even when the number of communities grows to infinity as the network size grows. Simulation study shows that the proposed method can recover true communities more accurately than other methods. Applications to a college football data and a yeast protein-protein interaction data also reveal that the proposed method performs significantly better.

stat.ME

Differential Network Analysis via the Lasso Penalized D-Trace Loss

Biological networks often change under different environmental and genetic conditions. Understanding how these networks change becomes an important problem in biological studies. In this paper, we model the network change as the difference of two precision matrices and propose a novel loss function for estimating the precision matrix difference. Under a new irrepresentability condition, we show that the new loss function with the lasso penalty can give consistent estimates in high-dimensional setting for sub-Gaussian and polynomial-tailed distributions. An efficient algorithm is developed based on the alternating direction method to solve the optimization problem. Simulation studies and a real data analysis about colorectal cancer show that the proposed method outperforms other available methods.

stat.ME