Searcharxiv⌕ Search

arXiv subjects

Ramon Diaz-Uriarte

Publications and source records attributed to Ramon Diaz-Uriarte.

9 recordsLinked to original sources

A method for comparing inferred evolutionary accumulation dynamics across covariates and model structures

In this note we describe a method for comparing inferred dynamics in evolutionary accumulation models (EvAMs). These models involve the acquisition of multiple, potentially codependent, binary features over time -- for example, mutations in cancer development, or phenotypes in evolutionary biology. As the set of methods for inferring EvAM dynamics expands, approaches for comparing inference across algorithms, datasets, and covariates are required. In particular, a comparison method supporting reversible, stochastic dynamics, interactions between feature sets, and potentially non-independent samples (as well as simpler cases) has yet to be established. It is possible for EvAM with similar relative feature orderings to produce completely different sets of observed states due to ``frameshift''-like differences; methods distinguishing state and transition similarity are therefore also desirable. Here we suggest a method focussed on ordering matrices, describing the probability that a feature is acquired under different conditions on other features, and is thus generally comparable across methods and datasets. We demonstrate how the approach captures both statistical robustness (given EvAM uncertainty) and scientifically meaningful differences in inferred dynamics. The method is applied to synthetic and real-world data on the evolution of chromosomal aberrations in different tumour types and of drug-resistant bacteria in different countries.

q-bio.QM↗

A structural causal framework for interventions on evolutionary accumulation models

Evolutionary accumulation models (EvAMs), also known as cancer progression models (CPMs), infer dependencies in the order of accumulation of mutations during tumor progression from cross-sectional data. It has been suggested that EvAMs could be used to identify therapeutic targets, but there is no procedure in the literature for how to extract predictions under intervention from these models. A simple approach of conditioning on the absence of a mutation gives incorrect predictions. We address this gap by formalizing what "intervene" means for all currently available EvAM methods (OT, OncoBN, CBN, H-ESBCN, MHN, HyperHMM, HyperTraPS), using Pearl's do operator and conditional interventions. For each model, we show how to implement the intervention (in most cases as specific parameter modifications), identify equivalent implementation procedures, and analyze whether the modularity assumption -- required for the intervention to be well-defined -- is justified. Drawing on individual-level causal DAGs that make fitness an explicit variable, we distinguish two types of intervention (killing and inactivating) that are conflated in standard EvAM representations. Since the goal is to prioritize intervention candidates, we recast the problem as one of ranking: we define three intervention objectives and provide a protocol for evaluating how well EvAMs rank targets. Our framework is not specific to cancer or EvAMs; it applies wherever fitted computational models can be interpreted as structural causal models. Code available from https://github.com/rdiaz02/scm-interv-evams.

q-bio.QM↗

Evolutionary accumulation modelling in AMR: machine learning to infer and predict evolutionary dynamics of multi-drug resistance

Can we understand and predict the evolutionary pathways by which bacteria acquire multi-drug resistance (MDR)? These questions have substantial potential impact in basic biology and in applied approaches to address the global health challenge of antimicrobial resistance (AMR). Here, we review how a class of machine learning approaches called evolutionary accumulation modelling (EvAM) may help reveal these dynamics using genetic and/or phenotypic AMR datasets, without requiring longitudinal sampling. These approaches are well-established in cancer progression and evolutionary biology, but currently less used in AMR research. We discuss how EvAM can learn the evolutionary pathways by which drug resistances and AMR features are acquired as pathogens evolve, predict next evolutionary steps, identify influences between AMR features, and explore differences in MDR evolution between regions, demographics, and more. We demonstrate a case study on MDR evolution in Mycobacterium tuberculosis and discuss the strengths and weaknesses of these approaches, providing links to some approaches for implementation.

q-bio.PE↗

A picture guide to cancer progression and monotonic accumulation models: evolutionary assumptions, plausible interpretations, and alternative uses

Cancer progression and monotonic accumulation models were developed to discover dependencies in the irreversible acquisition of binary traits from cross-sectional data. They have been used in computational oncology and virology but also in widely different problems such as malaria progression. These methods have been applied to predict future states of the system, identify routes of feature acquisition, and improve patient stratification, and they hold promise for evolutionary-based treatments. New methods continue to be developed. But these methods have shortcomings, which are yet to be systematically critiqued, regarding key evolutionary assumptions and interpretations. After an overview of the available methods, we focus on why inferences might not be about the processes we intend. Using fitness landscapes, we highlight difficulties that arise from bulk sequencing and reciprocal sign epistasis, from conflating lines of descent, path of the maximum, and mutational profiles, and from ambiguous use of the idea of exclusivity. We examine how the previous concerns change when bulk sequencing is explicitly considered, and underline opportunities for addressing dependencies due to frequency-dependent selection. This review identifies major standing issues, and should encourage the use of these methods in other areas with a better alignment between entities and model assumptions.

q-bio.PE↗

Global epistasis on fitness landscapes

Epistatic interactions between mutations add substantial complexity to adaptive landscapes, and are often thought of as detrimental to our ability to predict evolution. Yet, patterns of global epistasis, in which the fitness effect of a mutation is well-predicted by the fitness of its genetic background, may actually be of help in our efforts to reconstruct fitness landscapes and infer adaptive trajectories. Microscopic interactions between mutations, or inherent nonlinearities in the fitness landscape, may cause global epistasis patterns to emerge. In this brief review, we provide a succinct overview of recent work about global epistasis, with an emphasis on building intuition about why it is often observed. To this end, we reconcile simple geometric reasoning with recent mathematical analyses, using these to explain why different mutations in an empirical landscape may exhibit different global epistasis patterns - ranging from diminishing to increasing returns. Finally, we highlight open questions and research directions.

q-bio.PE↗

From genotypes to organisms: State-of-the-art and perspectives of a cornerstone in evolutionary dynamics

Understanding how genotypes map onto phenotypes, fitness, and eventually organisms is arguably the next major missing piece in a fully predictive theory of evolution. We refer to this generally as the problem of the genotype-phenotype map. Though we are still far from achieving a complete picture of these relationships, our current understanding of simpler questions, such as the structure induced in the space of genotypes by sequences mapped to molecular structures, has revealed important facts that deeply affect the dynamical description of evolutionary processes. Empirical evidence supporting the fundamental relevance of features such as phenotypic bias is mounting as well, while the synthesis of conceptual and experimental progress leads to questioning current assumptions on the nature of evolutionary dynamics-cancer progression models or synthetic biology approaches being notable examples. This work delves into a critical and constructive attitude in our current knowledge of how genotypes map onto molecular phenotypes and organismal functions, and discusses theoretical and empirical avenues to broaden and improve this comprehension. As a final goal, this community should aim at deriving an updated picture of evolutionary processes soundly relying on the structural properties of genotype spaces, as revealed by modern techniques of molecular and functional analysis.

q-bio.PE↗

Asterias: a parallelized web-based suite for the analysis of expression and aCGH data

Asterias (\url{http://www.asterias.info}) is an integrated collection of freely-accessible web tools for the analysis of gene expression and aCGH data. Most of the tools use parallel computing (via MPI). Most of our applications allow the user to obtain additional information for user-selected genes by using clickable links in tables and/or figures. Our tools include: normalization of expression and aCGH data; converting between different types of gene/clone and protein identifiers; filtering and imputation; finding differentially expressed genes related to patient class and survival data; searching for models of class prediction; using random forests to search for minimal models for class prediction or for large subsets of genes with predictive capacity; searching for molecular signatures and predictive genes with survival data; detecting regions of genomic DNA gain or loss. The capability to send results between different applications, access to additional functional information, and parallelized computation make our suite unique and exploit features only available to web-based applications.

q-bio.GN↗

Variable selection from random forests: application to gene expression data

Random forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations, and returns measures of variable importance. Thus, it is important to understand the performance of random forest with microarray data and its use for gene selection. We first show the effects of changes in parameters of random forest on the prediction error. Then we present an approach for gene selection that uses measures of variable importance and error rate, and is targeted towards the selection of small sets of genes. Using simulated and real microarray data, we show that the gene selection procedure yields small sets of genes while preserving predictive accuracy. Availability: All code is available as an R package, varSelRF, from CRAN, http://cran.r-project.org/src/contrib/PACKAGES.html, or from the supplementary material page. Supplementary information: http://ligarto.org/rdiaz/Papers/rfVS/randomForestVarSel.html

q-bio.QM↗

Molecular Signatures from Gene Expression Data

Motivation: ``Molecular signatures'' or ``gene-expression signatures'' are used to predict patients' characteristics using data from coexpressed genes. Signatures can enhance understanding about biological mechanisms and have diagnostic use. However, available methods to search for signatures fail to address key requirements of signatures, especially the discovery of sets of tightly coexpressed genes. Results: After suggesting an operational definition of signature, we develop a method that fulfills these requirements, returning sets of tightly coexpressed genes with good predictive performance. This method can also identify when the data are inconsistent with the hypothesis of a few, stable, easily interpretable sets of coexpressed genes. Identification of molecular signatures in some widely used data sets is questionable under this simple model, which emphasizes the needed for further work on the operationalization of the biological model and the assessment of the stability of putative signatures. Availability: The code (R with C++) is available from http://www.ligarto.org/rdiaz/Software/Software.html under the GNU GPL.

q-bio.QM↗