SearcharxivSearch

arXiv subjects

Andrea Califano

Publications and source records attributed to Andrea Califano.

11 recordsLinked to original sources

A Roadmap for Predictive Human Immunology

For over a century, immunology has masterfully discovered and dissected the components of our immune system, yet its collective behavior remains fundamentally unpredictable. In this perspective, we argue that building on the learnings of reductionist biology and systems immunology, the field is poised for a third revolution. This new era will be driven by the convergence of purpose-built, large-scale causal experiments and predictive, generalizable AI models. Here, we propose the Predictive Immunology Loop as the unifying engine to harness this convergence. This closed loop iteratively uses AI to design maximally informative experiments and, in turn, leverages the resulting data to improve dynamic, in silico models of the human immune system across biological scales, culminating in a Virtual Immune System. This engine provides a natural roadmap for addressing immunology's grand challenges, from decoding molecular recognition to engineering tissue ecosystems. It also offers a framework to transform immunology from a descriptive discipline into one capable of forecasting and, ultimately, engineering human health.

q-bio.OT

A path towards AI-scale, interoperable biological data

Biology is at the precipice of a new era where AI accelerates and amplifies the ability to study how cells operate, organize, and work as systems, revealing why disease happens and how to correct it. Organizations globally are prioritizing AI to accelerate basic research, drug discovery, personalized medicine, and synthetic biology. However, despite these opportunities, scientific data have proven a bottleneck, and progress has been slow and fragmented. Unless the scientific community takes a technology-led, community-focused approach to scaling and harnessing data, we will fail to capture this opportunity to drive new insights and biological discovery. The data bottleneck presents a unique paradox. It is increasingly simple to generate huge data volumes, thanks to expanding imaging datasets and plummeting sequencing costs, but scientists lack standards and tooling for large biological datasets, preventing integration into a multimodal foundational dataset that unlocks generalizable models of cellular and tissue function. This contradiction highlights two interrelated problems: abundant data that's difficult to manage, and a lack of data resources with necessary quality and utility to realize AI's potential in biology. Science must forge a collective approach enabling distributed contributions to combine into cohesive, powerful datasets transcending individual purposes. Here, we present a technological and data generation roadmap for scaling scientific impact. We outline AI's opportunity, mechanisms to scale data generation, the need for multi-modal measurements, and means to pool resources, standardize approaches, and collectively build the foundation enabling AI's full potential in biological discovery.

q-bio.OT

How to Build the Virtual Cell with Artificial Intelligence: Priorities and Opportunities

The cell is arguably the most fundamental unit of life and is central to understanding biology. Accurate modeling of cells is important for this understanding as well as for determining the root causes of disease. Recent advances in artificial intelligence (AI), combined with the ability to generate large-scale experimental data, present novel opportunities to model cells. Here we propose a vision of leveraging advances in AI to construct virtual cells, high-fidelity simulations of cells and cellular systems under different conditions that are directly learned from biological data across measurements and scales. We discuss desired capabilities of such AI Virtual Cells, including generating universal representations of biological entities across scales, and facilitating interpretable in silico experiments to predict and understand their behavior using virtual instruments. We further address the challenges, opportunities and requirements to realize this vision including data needs, evaluation strategies, and community standards and engagement to ensure biological accuracy and broad utility. We envision a future where AI Virtual Cells help identify new drug targets, predict cellular responses to perturbations, as well as scale hypothesis exploration. With open science collaborations across the biomedical ecosystem that includes academia, philanthropy, and the biopharma and AI industries, a comprehensive predictive understanding of cell mechanisms and interactions has come into reach.

q-bio.QM

High-throughput validation of ceRNA regulatory networks

Background: MicroRNAs (miRNAs) play multiple roles in tumor biology [1]. Interestingly, reports from multiple groups suggest that miRNA targets may be coupled through competitive stoichiometric sequestration [2]. Specifically, computational models predicted [3, 4] and experimental assays confirmed [5, 6] that miRNA activity is dependent on miRNA target abundance, and consequently, changes to the abundance of some miRNA targets lead to changes to the regulation and abundance of their other targets. The resulting indirect regulatory influence between miRNA targets resembles competition and has been dubbed competitive endogenous RNA (ceRNA) [5, 7, 8]. Recent studies have questioned the physiological relevance of ceRNA interactions [9], researchers ability to accurately predict these interactions [10], and the number of genes that are impacted by ceRNA interactions in specific cellular contexts [11]. Results: To address these concerns, we reverse engineered ceRNA networks (ceRNETs) in breast and prostate adenocarcinomas using context-specific TCGA profiles [12-14], and tested whether ceRNA interactions can predict the effects of RNAi-mediated gene silencing perturbations in PC3 and MCF7 cells. Our results, based on tests of thousands of inferred ceRNA interactions that are predicted to alter hundreds of cancer genes in each of the two tumor contexts, confirmed statistically significant effects for half of the predicted targets. Conclusions: Our results suggest that the expression of a significant fraction of cancer genes may be regulated by ceRNA interactions in each of the two tumor contexts.

q-bio.MN

Searching for Gene Sets with Mutually Exclusive Mutations

Cancer cells evolve through random somatic mutations. "Beneficial" mutations which disrupt key pathways (e.g. cell cycle regulation) are subject to natural selection. Multiple mutations may lead to the same "beneficial" effect, in which case there is no selective advantage to having more than one of these mutations. Hence we are interested in finding sets of genes whose mutations are approximately mutually exclusive (anti-co-occurring) within the TCGA Pancancer dataset. In principle, finding the best set is NP Hard. Nevertheless, we will show how a new Mutation anti-co-OCcurrence Algorithm (MOCA) provides an effective greedy search and testing algorithm with guaranteed control of the familywise error rate or false discovery rate, by combining some under-appreciated ideas from frequentist hypothesis testing. These ideas include: (a) A novel exact conditional test for the tendency of multiple sets to have a large/small union/intersection, which generalises Fisher's exact test of 2x2 tables. (b) Randomised hypothesis tests for discrete distributions. (c) Stouffer's method for combining p-values. (d) Weighted multiple hypothesis testing. A new approach to setting a-priori weights which generates additional implicit hypothesis tests is suggested, and allows us to preserve almost all statistical power when testing pairs despite introducing a combinatorially large number of additional hypotheses.

stat.ME

Multivariate dependence and genetic networks inference

A critical task in systems biology is the identification of genes that interact to control cellular processes by transcriptional activation of a set of target genes. Many methods have been developed to use statistical correlations in high-throughput datasets to infer such interactions. However, cellular pathways are highly cooperative, often requiring the joint effect of many molecules, and few methods have been proposed to explicitly identify such higher-order interactions, partially due to the fact that the notion of multivariate statistical dependency itself remains imprecisely defined. We define the concept of dependence among multiple variables using maximum entropy techniques and introduce computational tests for their identification. Synthetic network results reveal that this procedure uncovers dependencies even in undersampled regimes, when the joint probability distribution cannot be reliably estimated. Analysis of microarray data from human B cells reveals that third-order statistics, but not second-order ones, uncover relationships between genes that interact in a pathway to cooperatively regulate a common set of targets.

q-bio.QM

A Pattern Discovery-Based Method for Detecting Multi-Locus Genetic Association

Methods to effectively detect multi-locus genetic association are becoming increasingly relevant in the genetic dissection of complex trait in humans. Current approaches typically consider a limited number of hypotheses, most of which are related to the effect of a single locus or of a relatively small number of neighboring loci on a chromosomal region. We have developed a novel method that is specifically designed to detect genetic association involving multiple disease-susceptibility loci, possibly on different chromosomes. Our approach relies on the efficient discovery of patterns comprising spatially unrestricted polymorphic markers and on the use of appropriate test statistics to evaluate pattern-trait association. Power calculations using multi-locus disease models demonstrate significant gain of power by using this method in detecting multi-locus genetic association when compared to a standard single marker analysis method. When analyzing a Schizophrenia dataset, we confirmed a previously identified gene-gene interaction. In addition, a less conspicuous association involving different markers on the same two genes was also identified, implicating genetic heterogeneity.

q-bio.GN

Genome-wide discovery of modulators of transcriptional interactions in human B lymphocytes

Transcriptional interactions in a cell are modulated by a variety of mechanisms that prevent their representation as pure pairwise interactions between a transcription factor and its target(s). These include, among others, transcription factor activation by phosphorylation and acetylation, formation of active complexes with one or more co-factors, and mRNA/protein degradation and stabilization processes. This paper presents a first step towards the systematic, genome-wide computational inference of genes that modulate the interactions of specific transcription factors at the post-transcriptional level. The method uses a statistical test based on changes in the mutual information between a transcription factor and each of its candidate targets, conditional on the expression of a third gene. The approach was first validated on a synthetic network model, and then tested in the context of a mammalian cellular system. By analyzing 254 microarray expression profiles of normal and tumor related human B lymphocytes, we investigated the post transcriptional modulators of the MYC proto-oncogene, an important transcription factor involved in tumorigenesis. Our method discovered a set of 100 putative modulator genes, responsible for modulating 205 regulatory relationships between MYC and its targets. The set is significantly enriched in molecules with function consistent with their activities as modulators of cellular interactions, recapitulates established MYC regulation pathways, and provides a notable repertoire of novel regulators of MYC function. The approach has broad applicability and can be used to discover modulators of any other transcription factor, provided that adequate expression profile data are available.

q-bio.MN

ARACNE: An Algorithm for the Reconstruction of Gene Regulatory Networks in a Mammalian Cellular Context

Background: Elucidating gene regulatory networks is crucial for understanding normal cell physiology and complex pathologic phenotypes. Existing computational methods for the genome-wide ``reverse engineering'' of such networks have been successful only for lower eukaryotes with simple genomes. Here we present ARACNE, a novel algorithm, using microarray expression profiles, specifically designed to scale up to the complexity of regulatory networks in mammalian cells, yet general enough to address a wider range of network deconvolution problems. This method uses an information theoretic approach to eliminate the majority of indirect interactions inferred by co-expression methods. Results: We prove that ARACNE reconstructs the network exactly (asymptotically) if the effect of loops in the network topology is negligible, and we show that the algorithm works well in practice, even in the presence of numerous loops and complex topologies. We assess ARACNE's ability to reconstruct transcriptional regulatory networks using both a realistic synthetic dataset and a microarray dataset from human B cells. On synthetic datasets ARACNE achieves very low error rates and outperforms established methods, such as Relevance Networks and Bayesian Networks. Application to the deconvolution of genetic networks in human B cells demonstrates ARACNE's ability to infer validated transcriptional targets of the c MYC proto-oncogene. We also study the effects of mis estimation of mutual information on network reconstruction, and show that algorithms based on mutual information ranking are more resilient to estimation errors.

q-bio.MN

Conditional Network Analysis Identifies Candidate Regulator Genes in Human B Cells

Cellular phenotypes are determined by the dynamical activity of networks of co-regulated genes. Elucidating such networks is crucial for the understanding of normal cell physiology as well as for the dissection of complex pathologic phenotypes. Existing methods for such "reverse engineering" of genetic networks from microarray expression data have been successful only in prokaryotes (E. coli) and lower eukaryotes (S. cerevisiae) with relatively simple genomes. Additionally, they have mostly attempted to reconstruct average properties about the network connectivity without capturing the highly conditional nature of the interactions. In this paper we extend the ARACNE algorithm, which we recently introduced and successfully applied to the reconstruction of whole-genome transcriptional networks from mammalian cells, precisely to link the existence of specific network structures to the expression or lack thereof of specific regulator genes. This is accomplished by analyzing thousands of alternative network topologies generated by constraining the data set on the presence or absence of putative regulator genes. By considering interactions that are consistently supported across several such constraints, we identify many transcriptional interactions that would not have been detectable by the original method. By selecting genes that produce statistically significant changes in network topology, we identify novel candidate regulator genes. Further analysis shows that transcription factors, kinases, phosphatases, and other gene families known to effect biochemical interactions, are significantly overrepresented among the set of candidate regulator genes identified in silico, indirectly supporting the validity of the approach.

q-bio.MN