SearcharxivSearch

arXiv subjects

Lingxue Zhu

Publications and source records attributed to Lingxue Zhu.

6 recordsLinked to original sources

A Unified Statistical Framework for Single Cell and Bulk RNA Sequencing Data

Recent advances in technology have enabled the measurement of RNA levels for individual cells. Compared to traditional tissue-level bulk RNA-seq data, single cell sequencing yields valuable insights about gene expression profiles for different cell types, which is potentially critical for understanding many complex human diseases. However, developing quantitative tools for such data remains challenging because of high levels of technical noise, especially the "dropout" events. A "dropout" happens when the RNA for a gene fails to be amplified prior to sequencing, producing a "false" zero in the observed data. In this paper, we propose a Unified RNA-Sequencing Model (URSM) for both single cell and bulk RNA-seq data, formulated as a hierarchical model. URSM borrows the strength from both data sources and carefully models the dropouts in single cell data, leading to a more accurate estimation of cell type specific gene expression profile. In addition, URSM naturally provides inference on the dropout entries in single cell data that need to be imputed for downstream analyses, as well as the mixing proportions of different cell types in bulk samples. We adopt an empirical Bayes approach, where parameters are estimated using the EM algorithm and approximate inference is obtained by Gibbs sampling. Simulation results illustrate that URSM outperforms existing approaches both in correcting for dropouts in single cell data, as well as in deconvolving bulk samples. We also demonstrate an application to gene expression data on fetal brains, where our model successfully imputes the dropout genes and reveals cell type specific expression patterns.

stat.AP

Deep and Confident Prediction for Time Series at Uber

Reliable uncertainty estimation for time series prediction is critical in many fields, including physics, biology, and manufacturing. At Uber, probabilistic time series forecasting is used for robust prediction of number of trips during special events, driver incentive allocation, as well as real-time anomaly detection across millions of metrics. Classical time series models are often used in conjunction with a probabilistic formulation for uncertainty estimation. However, such models are hard to tune, scale, and add exogenous variables to. Motivated by the recent resurgence of Long Short Term Memory networks, we propose a novel end-to-end Bayesian deep model that provides time series prediction along with uncertainty estimation. We provide detailed experiments of the proposed solution on completed trips data, and successfully apply it to large-scale time series anomaly detection at Uber.

stat.ML

Testing High Dimensional Covariance Matrices, with Application to Detecting Schizophrenia Risk Genes

Scientists routinely compare gene expression levels in cases versus controls in part to determine genes associated with a disease. Similarly, detecting case-control differences in co-expression among genes can be critical to understanding complex human diseases; however statistical methods have been limited by the high dimensional nature of this problem. In this paper, we construct a sparse-Leading-Eigenvalue-Driven (sLED) test for comparing two high-dimensional covariance matrices. By focusing on the spectrum of the differential matrix, sLED provides a novel perspective that accommodates what we assume to be common, namely sparse and weak signals in gene expression data, and it is closely related with Sparse Principal Component Analysis. We prove that sLED achieves full power asymptotically under mild assumptions, and simulation studies verify that it outperforms other existing procedures under many biologically plausible scenarios. Applying sLED to the largest gene-expression dataset obtained from post-mortem brain tissue from Schizophrenia patients and controls, we provide a novel list of genes implicated in Schizophrenia and reveal intriguing patterns in gene co-expression change for Schizophrenia subjects. We also illustrate that sLED can be generalized to compare other gene-gene "relationship" matrices that are of practical interest, such as the weighted adjacency matrices.

stat.ME

A Generic Sample Splitting Approach for Refined Community Recovery in Stochastic Block Models

We propose and analyze a generic method for community recovery in stochastic block models and degree corrected block models. This approach can exactly recover the hidden communities with high probability when the expected node degrees are of order $\log n$ or higher. Starting from a roughly correct community partition given by some conventional community recovery algorithm, this method refines the partition in a cross clustering step. Our results simplify and extend some of the previous work on exact community recovery, discovering the key role played by sample splitting. The proposed method is simple and can be implemented with many practical community recovery algorithms.

stat.ML

Continuous Interior Penalty Finite Element Method for Helmholtz Equation with High Wave Number: One Dimensional Analysis

This paper addresses the properties of Continuous Interior Penalty (CIP) finite element solutions for the Helmholtz equation. The $h$-version of the CIP finite element method with piecewise linear approximation is applied to a one-dimensional model problem. We first show discrete well posedness and convergence results, using the imaginary part of the stabilization operator, for the complex Helmholtz equation. Then we consider a method with real valued penalty parameter and prove an error estimate of the discrete solution in the $H^1$-norm, as the sum of best approximation plus a pollution term that is the order of the phase difference. It is proved that the pollution can be eliminated by selecting the penalty parameter appropriately. As a result of this analysis, thorough and rigorous understanding of the error behavior throughout the range of convergence is gained. Numerical results are presented that show sharpness of the error estimates and highlight some phenomena of the discrete solution behavior.

math.NA

Pre-asymptotic Error Analysis of CIP-FEM and FEM for Helmholtz Equation with High Wave Number. Part II: $hp$ version

In this paper, which is part II in a series of two, the pre-asymptotic error analysis of the continuous interior penalty finite element method (CIP-FEM) and the FEM for the Helmholtz equation in two and three dimensions is continued. While part I contained results on the linear CIP-FEM and FEM, the present part deals with approximation spaces of order $p \ge 1$. By using a modified duality argument, pre-asymptotic error estimates are derived for both methods under the condition of $\frac{kh}{p}\le C_0\big(\frac{p}{k}\big)^{\frac{1}{p+1}}$, where $k$ is the wave number, $h$ is the mesh size, and $C_0$ is a constant independent of $k, h, p$, and the penalty parameters. It is shown that the pollution errors of both methods in $H^1$-norm are $O(k^{2p+1}h^{2p})$ if $p=O(1)$ and are $O\Big(\frac{k}{p^2}\big(\frac{kh}{σp}\big)^{2p}\Big)$ if the exact solution $u\in H^2(\Om)$ which coincide with existent dispersion analyses for the FEM on Cartesian grids. Here $\si$ is a constant independent of $k, h, p$, and the penalty parameters. Moreover, it is proved that the CIP-FEM is stable for any $k, h, p>0$ and penalty parameters with positive imaginary parts. Besides the advantage of the absolute stability of the CIP-FEM compared to the FEM, the penalty parameters may be tuned to reduce the pollution effects.

math.NA