SearcharxivSearch

arXiv subjects

Nikolai Slavov

Publications and source records attributed to Nikolai Slavov.

15 recordsLinked to original sources

Single-Cell Proteomic Technologies: Tools in the quest for principles

Over the last decade, proteomic analysis of single cells by mass spectrometry transitioned from an uncertain possibility to a set of robust and rapidly advancing technologies supporting the accurate quantification of thousands of proteins. We review the major drivers of this progress, from establishing feasibility to powerful and increasingly scalable methods. We focus on the tradeoffs and synergies of different technological solutions within a coherent conceptual framework, which projects considerable room both for throughput scaling and for extending the analysis scope to functional protein measurements. We highlight the potential of these technologies to support the development of mechanistic biophysical models and help uncover new principles.

q-bio.QM

Modeling and interpretation of single-cell proteogenomic data

Biological functions stem from coordinated interactions among proteins, nucleic acids and small molecules. Mass spectrometry technologies for reliable, high throughput single-cell proteomics will add a new modality to genomics and enable data-driven modeling of the molecular mechanisms coordinating proteins and nucleic acids at single-cell resolution. This promising potential requires estimating the reliability of measurements and computational analysis so that models can distinguish biological regulation from technical artifacts. We highlight different measurement modes that can support single-cell proteogenomic analysis and how to estimate their reliability. We then discuss approaches for developing both abstract and mechanistic models that aim to biologically interpret the measured differences across modalities, including specific applications to directed stem cell differentiation and to inferring protein interactions in cancer cells from the buffing of DNA copy-number variations. Single-cell proteogenomic data will support mechanistic models of direct molecular interactions that will provide generalizable and predictive representations of biological systems.

q-bio.GN

Sampling the proteome by emerging single-molecule and mass-spectrometry methods

Mammalian cells have about 30,000-fold more protein molecules than mRNA molecules. This larger number of molecules and the associated larger dynamic range have major implications in the development of proteomics technologies. We examine these implications for both liquid chromatography-tandem mass spectrometry (LC-MS/MS) and single-molecule counting and provide estimates on how many molecules are routinely measured in proteomics experiments by LC-MS/MS. We review strategies that have been helpful for counting billions of protein molecules by LC-MS/MS and suggest that these strategies can benefit single-molecule methods, especially in mitigating the challenges of the wide dynamic range of the proteome. We also examine the theoretical possibilities for scaling up single-molecule and mass spectrometry proteomics approaches to quantifying the billions of protein molecules that make up the proteomes of our cells.

q-bio.QM

Initial recommendations for performing, benchmarking, and reporting single-cell proteomics experiments

Analyzing proteins from single cells by tandem mass spectrometry (MS) has become technically feasible. While such analysis has the potential to accurately quantify thousands of proteins across thousands of single cells, the accuracy and reproducibility of the results may be undermined by numerous factors affecting experimental design, sample preparation, data acquisition, and data analysis. Broadly accepted community guidelines and standardized metrics will enhance rigor, data quality, and alignment between laboratories. Here we propose best practices, quality controls, and data reporting recommendations to assist in the broad adoption of reliable quantitative workflows for single-cell proteomics.

q-bio.OT

New views of old proteins: clarifying the enigmatic proteome

All human diseases involve proteins, yet our current tools to characterize and quantify them are limited. To better elucidate proteins across space, time, and molecular composition, we provide provocative projections for technologies to meet the challenges that protein biology presents. With a broad perspective, we discuss grand opportunities to transition the science of proteomics into a more propulsive enterprise. Extrapolating recent trends, we offer potential futures for a next generation of disruptive approaches to define, quantify and visualize the multiple dimensions of the proteome, thereby transforming our understanding and interactions with human disease in the coming decade.

q-bio.BM

Analyzing ribosome remodeling in health and disease

regulation largely unexplored, in part due to methodological limitations. Indeed, we review evidence demonstrating that commonly used methods, such as transcriptomics, are inadequate because the variability in mRNAs coding for ribosomal proteins (RP) does not necessarily correspond to RP variability. Thus protein remodeling of ribosomes should be investigated by methods that allow direct quantification of RPs, ideally of isolated ribosomes. We review such methods, focusing on mass spectrometry and emphasizing method-specific biases and approaches to control these biases. We argue that using multiple complementary methods can help reduce the danger of interpreting reproducible systematic biases as evidence for ribosome remodeling.

q-bio.QM

Single-cell protein analysis by mass-spectrometry

Human physiology and pathology arise from the coordinated interactions of diverse single cells. However, analyzing single cells has been limited by the low sensitivity and throughput of analytical methods. DNA sequencing has recently made such analysis feasible for nucleic acids, but single-cell protein analysis remains limited. Mass-spectrometry is the most powerful method for protein analysis, but its application to single cells faces three major challenges: Efficiently delivering proteins/peptides to MS detectors, identifying their sequences, and scaling the analysis to many thousands of single cells. These challenges have motivated corresponding solutions, including SCoPE-design multiplexing and clean, automated, and miniaturized sample preparation. Synergistically applied, these solutions enable quantifying thousands of proteins across many single cells and establish a solid foundation for further advances. Building upon this foundation, the SCoPE concept will enable analyzing subcellular organelles and post-translational modifications while increases in multiplexing capabilities will increase the throughput and decrease cost.

q-bio.QM

Mass-spectrometry of single mammalian cells quantifies proteome heterogeneity during cell differentiation

Cellular heterogeneity is important to biological processes, including cancer and development. However, proteome heterogeneity is largely unexplored because of the limitations of existing methods for quantifying protein levels in single cells. To alleviate these limitations, we developed Single Cell ProtEomics by Mass Spectrometry (SCoPE-MS), and validated its ability to identify distinct human cancer cell types based on their proteomes. We used SCoPE-MS to quantify over a thousand proteins in differentiating mouse embryonic stem (ES) cells. The single-cell proteomes enabled us to deconstruct cell populations and infer protein abundance relationships. Comparison between single-cell proteomes and transcriptomes indicated coordinated mRNA and protein covariation. Yet many genes exhibited functionally concerted and distinct regulatory patterns at the mRNA and the protein levels, suggesting that post-transcriptional regulatory mechanisms contribute to proteome remodeling during lineage specification, especially for developmental genes. SCoPE-MS is broadly applicable to measuring proteome configurations of single cells and linking them to functional phenotypes, such as cell type and differentiation potentials.

q-bio.GN

Routinely quantifying single cell proteomes: A new age in quantitative biology and medicine

Many pressing medical challenges - such as diagnosing disease, enhancing directed stem cell differentiation, and classifying cancers - have long been hindered by limitations in our ability to quantify proteins in single cells. Mass-spectrometry (MS) is poised to transcend these limitations by developing powerful methods to routinely quantify thousands of proteins and proteoforms across many thousands of single cells. We outline specific technological developments and ideas that can increase the sensitivity and throughput of single cell MS by orders of magnitude and usher in this new age. These advances will transform medicine and ultimately contribute to understanding biological systems on an entirely new level.

q-bio.GN

Quantifying homologous proteins and proteoforms

Many proteoforms - arising from alternative splicing, post-translational modifications (PTMs), or paralogous genes - have distinct biological functions, such as histone PTM proteoforms. However, their quantification by existing bottom-up mass-spectrometry (MS) methods is undermined by peptide-specific biases. To avoid these biases, we developed and implemented a first-principles model (HIquant) for quantifying proteoform stoichiometries. We characterized when MS data allow inferring proteoform stoichiometries by HIquant, derived an algorithm for optimal inference, and demonstrated experimentally high accuracy in quantifying fractional PTM occupancy without using external standards, even in the challenging case of the histone modification code. HIquant server is implemented at: https://web.northeastern.edu/slavov/2014_HIquant/

q-bio.QM

Post-transcriptional regulation across human tissues

Transcriptional and post-transcriptional regulation shape tissue-type-specific proteomes, but their relative contributions remain contested. Estimates of the factors determining protein levels in human tissues do not distinguish between (i) the factors determining the variability between the abundances of different proteins, i.e., mean-level-variability and, (ii) the factors determining the physiological variability of the same protein across different tissue types, i.e., across-tissues variability. We sought to estimate the contribution of transcript levels to these two orthogonal sources of variability, and found that scaled mRNA levels can account for most of the mean-level-variability but not necessarily for across-tissues variability. The reliable quantification of the latter estimate is limited by substantial measurement noise. However, protein-to-mRNA ratios exhibit substantial across-tissues variability that is functionally concerted and reproducible across different datasets, suggesting extensive post-transcriptional regulation. These results caution against estimating protein fold-changes from mRNA fold-changes between different cell-types, and highlight the contribution of post-transcriptional regulation to shaping tissue-type-specific proteomes.

q-bio.GN

Differential stoichiometry among core ribosomal proteins

Understanding the regulation and structure of ribosomes is essential to understanding protein synthesis and its deregulation in disease. While ribosomes are believed to have a fixed stoichiometry among their core ribosomal proteins (RPs), some experiments suggest a more variable composition. Testing such variability requires direct and precise quantification of RPs. We used mass-spectrometry to directly quantify RPs across monosomes and polysomes of mouse embryonic stem cells (ESC) and budding yeast. Our data show that the stoichiometry among core RPs in wild-type yeast cells and ESC depends both on the growth conditions and on the number of ribosomes bound per mRNA. Furthermore, we find that the fitness of cells with a deleted RP-gene is inversely proportional to the enrichment of the corresponding RP in polysomes. Together, our findings support the existence of ribosomes with distinct protein composition and physiological function.

q-bio.GN

Extensive regulation of metabolism and growth during the cell division cycle

Yeast cells grown in culture can spontaneously synchronize their respiration, metabolism, gene expression and cell division. Such metabolic oscillations in synchronized cultures reflect single-cell oscillations, but the relationship between the oscillations in single cells and synchronized cultures is poorly understood. To understand this relationship and the coordination between metabolism and cell division, we collected and analyzed DNA-content, gene-expression and physiological data, at hundreds of time-points, from cultures metabolically-synchronized at different growth rates, carbon sources and biomass densities. The data enabled us to extend and generalize an ensemble-average-over-phases (EAP) model that connects the population-average gene-expression of asynchronous cultures to the gene-expression dynamics in the single-cells comprising the cultures. The extended model explains the carbon-source specific growth-rate responses of hundreds of genes. Our data demonstrate that for a given growth rate, the frequency of metabolic cycling in synchronized cultures increases with the biomass density. This observation underscores the difference between metabolic cycling in synchronized cultures and in single cells and suggests entraining of the single-cell cycle by a quorum-sensing mechanism. Constant levels of residual glucose during the metabolic cycling of synchronized cultures indicate that storage carbohydrates are required to fuel not only the G1/S transition of the division cycle but also the metabolic cycle. Despite the large variation in profiled conditions and in the time-scale of their dynamics, most genes preserve invariant dynamics of coordination with each other and with the rate of oxygen consumption. Similarly, the G1/S transition always occurs at the beginning, middle or end of the high oxygen consumption phases, analogous to observations in human and drosophila cells.

q-bio.GN

Convex Total Least Squares

We study the total least squares (TLS) problem that generalizes least squares regression by allowing measurement errors in both dependent and independent variables. TLS is widely used in applied fields including computer vision, system identification and econometrics. The special case when all dependent and independent variables have the same level of uncorrelated Gaussian noise, known as ordinary TLS, can be solved by singular value decomposition (SVD). However, SVD cannot solve many important practical TLS problems with realistic noise structure, such as having varying measurement noise, known structure on the errors, or large outliers requiring robust error-norms. To solve such problems, we develop convex relaxation approaches for a general class of structured TLS (STLS). We show both theoretically and experimentally, that while the plain nuclear norm relaxation incurs large approximation errors for STLS, the re-weighted nuclear norm approach is very effective, and achieves better accuracy on challenging STLS problems than popular non-convex solvers. We describe a fast solution based on augmented Lagrangian formulation, and apply our approach to an important class of biological problems that use population average measurements to infer cell-type and physiological-state specific expression levels that are very hard to measure directly.

stat.ML

Inference of Sparse Networks with Unobserved Variables. Application to Gene Regulatory Networks

Networks are a unifying framework for modeling complex systems and network inference problems are frequently encountered in many fields. Here, I develop and apply a generative approach to network inference (RCweb) for the case when the network is sparse and the latent (not observed) variables affect the observed ones. From all possible factor analysis (FA) decompositions explaining the variance in the data, RCweb selects the FA decomposition that is consistent with a sparse underlying network. The sparsity constraint is imposed by a novel method that significantly outperforms (in terms of accuracy, robustness to noise, complexity scaling, and computational efficiency) Bayesian methods and MLE methods using l1 norm relaxation such as K-SVD and l1--based sparse principle component analysis (PCA). Results from simulated models demonstrate that RCweb recovers exactly the model structures for sparsity as low (as non-sparse) as 50% and with ratio of unobserved to observed variables as high as 2. RCweb is robust to noise, with gradual decrease in the parameter ranges as the noise level increases.

stat.ML