Searcharxiv⌕ Search

arXiv subjects

Iain G. Johnston

Publications and source records attributed to Iain G. Johnston.

At least 19 recordsLinked to original sources

A method for comparing inferred evolutionary accumulation dynamics across covariates and model structures

In this note we describe a method for comparing inferred dynamics in evolutionary accumulation models (EvAMs). These models involve the acquisition of multiple, potentially codependent, binary features over time -- for example, mutations in cancer development, or phenotypes in evolutionary biology. As the set of methods for inferring EvAM dynamics expands, approaches for comparing inference across algorithms, datasets, and covariates are required. In particular, a comparison method supporting reversible, stochastic dynamics, interactions between feature sets, and potentially non-independent samples (as well as simpler cases) has yet to be established. It is possible for EvAM with similar relative feature orderings to produce completely different sets of observed states due to ``frameshift''-like differences; methods distinguishing state and transition similarity are therefore also desirable. Here we suggest a method focussed on ordering matrices, describing the probability that a feature is acquired under different conditions on other features, and is thus generally comparable across methods and datasets. We demonstrate how the approach captures both statistical robustness (given EvAM uncertainty) and scientifically meaningful differences in inferred dynamics. The method is applied to synthetic and real-world data on the evolution of chromosomal aberrations in different tumour types and of drug-resistant bacteria in different countries.

q-bio.QM↗

A structural causal framework for interventions on evolutionary accumulation models

Evolutionary accumulation models (EvAMs), also known as cancer progression models (CPMs), infer dependencies in the order of accumulation of mutations during tumor progression from cross-sectional data. It has been suggested that EvAMs could be used to identify therapeutic targets, but there is no procedure in the literature for how to extract predictions under intervention from these models. A simple approach of conditioning on the absence of a mutation gives incorrect predictions. We address this gap by formalizing what "intervene" means for all currently available EvAM methods (OT, OncoBN, CBN, H-ESBCN, MHN, HyperHMM, HyperTraPS), using Pearl's do operator and conditional interventions. For each model, we show how to implement the intervention (in most cases as specific parameter modifications), identify equivalent implementation procedures, and analyze whether the modularity assumption -- required for the intervention to be well-defined -- is justified. Drawing on individual-level causal DAGs that make fitness an explicit variable, we distinguish two types of intervention (killing and inactivating) that are conflated in standard EvAM representations. Since the goal is to prioritize intervention candidates, we recast the problem as one of ranking: we define three intervention objectives and provide a protocol for evaluating how well EvAMs rank targets. Our framework is not specific to cancer or EvAMs; it applies wherever fitted computational models can be interpreted as structural causal models. Code available from https://github.com/rdiaz02/scm-interv-evams.

q-bio.QM↗

Flexible inference of evolutionary accumulation dynamics using uncertain observational data

Understanding and predicting evolutionary accumulation pathways is a key objective in many fields of research, ranging from classical evolutionary biology to diverse applications in medicine. In this context, we are often confronted with the problem that data is sparse and uncertain. To use the available data as best as possible, inference approaches that can handle this uncertainty are required. One way that allows us to use not only cross-sectional data, but also phylogenetic related and longitudinal data, is using `hypercubic inference' models. In this article we introduce HyperLAU, a new algorithm for hypercubic inference that makes it possible to use datasets including uncertainties for learning evolutionary pathways. Expanding the flexibility of accumulation modelling, HyperLAU allows us to infer dynamic pathways and interactions between features, even when large sets of particular features are unobserved across the source dataset. We show that HyperLAU is able to highlight the main pathways found by other tools, even when up to 50% of the features in the input data are uncertain. Additionally, we demonstrate how it can help to overcome possible biases that can occur then reducing the used data by excluding uncertain parts. We illustrate the approach with a case study on multidrug resistance in tuberculosis, showing that HyperLAU allows more flexible data and provides new information about evolutionary pathways compared to existing approaches.

q-bio.PE↗

An Algebraic Approach to Evolutionary Accumulation Models

We present an algebraic approach to evolutionary accumulation modelling (EvAM). EvAM is concerned with learning and predicting the order in which evolutionary features accumulate over time. Our approach is complementary to the more common optimisation-based inference methods used in this field. Namely, we first use the natural underlying polynomial structure of the evolutionary process to define a semi-algebraic set of candidate parameters consistent with a given data set before maximising the likelihood function. We consider explicit examples and show that this approach is compatible with the solutions given by various statistical evolutionary accumulation models. Furthermore, we discuss the additional information of our algebraic model relative to these models.

stat.AP↗

Extracting useful information about reversible evolutionary processes from irreversible evolutionary accumulation models

Evolutionary accumulation models (EvAMs) are an emerging class of machine learning methods designed to infer the evolutionary pathways by which features are acquired. Applications include cancer evolution (accumulation of mutations), anti-microbial resistance (accumulation of drug resistances), genome evolution (organelle gene transfers), and more diverse themes in biology and beyond. Following these themes, many EvAMs assume that features are gained irreversibly -- no loss of features can occur. Reversible approaches do exist but are often computationally (much) more demanding and statistically less stable. Our goal here is to explore whether useful information about evolutionary dynamics which are in reality reversible can be obtained from modelling approaches with an assumption of irreversibility. We identify, and use simulation studies to quantify, errors involved in neglecting reversible dynamics, and show the situations in which approximate results from tractable models can be informative and reliable. In particular, EvAM inferences about the relative orderings of acquisitions and the core dynamic structure of evolutionary pathways -- which features are likely present when another is acquired -- are robust to reversibility in many cases, while estimations of uncertainty and feature interactions are more error-prone.

q-bio.PE↗

Evolutionary accumulation modelling in AMR: machine learning to infer and predict evolutionary dynamics of multi-drug resistance

Can we understand and predict the evolutionary pathways by which bacteria acquire multi-drug resistance (MDR)? These questions have substantial potential impact in basic biology and in applied approaches to address the global health challenge of antimicrobial resistance (AMR). Here, we review how a class of machine learning approaches called evolutionary accumulation modelling (EvAM) may help reveal these dynamics using genetic and/or phenotypic AMR datasets, without requiring longitudinal sampling. These approaches are well-established in cancer progression and evolutionary biology, but currently less used in AMR research. We discuss how EvAM can learn the evolutionary pathways by which drug resistances and AMR features are acquired as pathogens evolve, predict next evolutionary steps, identify influences between AMR features, and explore differences in MDR evolution between regions, demographics, and more. We demonstrate a case study on MDR evolution in Mycobacterium tuberculosis and discuss the strengths and weaknesses of these approaches, providing links to some approaches for implementation.

q-bio.PE↗

A picture guide to cancer progression and monotonic accumulation models: evolutionary assumptions, plausible interpretations, and alternative uses

Cancer progression and monotonic accumulation models were developed to discover dependencies in the irreversible acquisition of binary traits from cross-sectional data. They have been used in computational oncology and virology but also in widely different problems such as malaria progression. These methods have been applied to predict future states of the system, identify routes of feature acquisition, and improve patient stratification, and they hold promise for evolutionary-based treatments. New methods continue to be developed. But these methods have shortcomings, which are yet to be systematically critiqued, regarding key evolutionary assumptions and interpretations. After an overview of the available methods, we focus on why inferences might not be about the processes we intend. Using fitness landscapes, we highlight difficulties that arise from bulk sequencing and reciprocal sign epistasis, from conflating lines of descent, path of the maximum, and mutational profiles, and from ambiguous use of the idea of exclusivity. We examine how the previous concerns change when bulk sequencing is explicitly considered, and underline opportunities for addressing dependencies due to frequency-dependent selection. This review identifies major standing issues, and should encourage the use of these methods in other areas with a better alignment between entities and model assumptions.

q-bio.PE↗

Encounter networks from collective mitochondrial dynamics support the emergence of effective mtDNA genomes in plant cells

Mitochondria in plant cells form strikingly dynamic populations of largely individual organelles. Each mitochondrion contains on average less than a full copy of the mitochondrial DNA (mtDNA) genome. Here, we asked whether mitochondrial dynamics may allow individual mitochondria to `collect' a full copy of the mtDNA genome over time, by facilitating exchange between individuals. Akin to trade on a social network, exchange of mtDNA fragments across organelles may lead to the emergence of full `effective' genomes in individuals over time. We characterise the collective dynamics of mitochondria in \emph{Arabidopsis thaliana} hypocotyl cells using a recent approach combining single-cell timelapse microscopy, video analysis, and network science. We then use a quantitative model to predict the capacity for the sharing and accumulation of genetic information through the networks of encounters between mitochondria. We find that biological encounter networks are strikingly well predisposed to support the collection of full genomes over time, outperforming a range of other networks generated from theory and simulation. Using results from the coupon collector's problem, we show that the upper tail of the degree distribution is a key determinant of an encounter network's performance at this task and discuss how features of mitochondrial dynamics observed in biology facilitate the emergence of full effective genomes.

q-bio.SC↗

Data-driven modelling and characterisation of task completion sequences in online courses

The intrinsic temporality of learning demands the adoption of methodologies capable of exploiting time-series information. In this study we leverage the sequence data framework and show how data-driven analysis of temporal sequences of task completion in online courses can be used to characterise personal and group learners' behaviors, and to identify critical tasks and course sessions in a given course design. We also introduce a recently developed probabilistic Bayesian model to learn sequence trajectories of students and predict student performance. The application of our data-driven sequence-based analyses to data from learners undertaking an on-line Business Management course reveals distinct behaviors within the cohort of learners, identifying learners or groups of learners that deviate from the nominal order expected in the course. Using course grades a posteriori, we explore differences in behavior between high and low performing learners. We find that high performing learners follow the progression between weekly sessions more regularly than low performing learners, yet within each weekly session high performing learners are less tied to the nominal task order. We then model the sequences of high and low performance students using the probablistic Bayesian model and show that we can learn engagement behaviors associated with performance. We also show that the data sequence framework can be used for task centric analysis; we identify critical junctures and differences among types of tasks within the course design. We find that non-rote learning tasks, such as interactive tasks or discussion posts, are correlated with higher performance. We discuss the application of such analytical techniques as an aid to course design, intervention, and student supervision.

cs.SI↗

Optimal strategies in the Fighting Fantasy gaming system: influencing stochastic dynamics by gambling with limited resource

Fighting Fantasy is a popular recreational fantasy gaming system worldwide. Combat in this system progresses through a stochastic game involving a series of rounds, each of which may be won or lost. Each round, a limited resource (`luck') may be spent on a gamble to amplify the benefit from a win or mitigate the deficit from a loss. However, the success of this gamble depends on the amount of remaining resource, and if the gamble is unsuccessful, benefits are reduced and deficits increased. Players thus dynamically choose to expend resource to attempt to influence the stochastic dynamics of the game, with diminishing probability of positive return. The identification of the optimal strategy for victory is a Markov decision problem that has not yet been solved. Here, we combine stochastic analysis and simulation with dynamic programming to characterise the dynamical behaviour of the system in the absence and presence of gambling policy. We derive a simple expression for the victory probability without luck-based strategy. We use a backward induction approach to solve the Bellman equation for the system and identify the optimal strategy for any given state during the game. The optimal control strategies can dramatically enhance success probabilities, but take detailed forms; we use stochastic simulation to approximate these optimal strategies with simple heuristics that can be practically employed. Our findings provide a roadmap to improving success in the games that millions of people play worldwide, and inform a class of resource allocation problems with diminishing returns in stochastic games.

cs.AI↗

Intracellular Energy Variability Modulates Cellular Decision-Making Capacity

Cells are able to generate phenotypic diversity both during development and in response to stressful and changing environments, aiding survival. The biologically and medically vital process of a cell assuming a functionally important fate from a range of phenotypic possibilities can be thought of as a cell decision. To make these decisions, a cell relies on energy dependent pathways of signalling and expression. However, energy availability is often overlooked as a modulator of cellular decision-making. As cells can vary dramatically in energy availability, this limits our knowledge of how this key biological axis affects cell behaviour. Here, we consider the energy dependence of a highly generalisable decision-making regulatory network, and show that energy variability changes the sets of decisions a cell can make and the ease with which they can be made. Increasing intracellular energy levels can increase the number of stable phenotypes it can generate, corresponding to increased decision-making capacity. For this decision-making architecture, a cell with intracellular energy below a threshold is limited to a singular phenotype, potentially forcing the adoption of a specific cell fate. We suggest that common energetic differences between cells may explain some of the observed variability in cellular decision-making, and demonstrate the importance of considering energy levels in several diverse biological decision-making phenomena.

q-bio.CB↗

HyperTraPS: Inferring probabilistic patterns of trait acquisition in evolutionary and disease progression pathways

The explosion of data throughout the biomedical sciences provides unprecedented opportunities to learn about the dynamics of evolution and disease progression, but harnessing these large and diverse datasets remains challenging. Here, we describe a highly generalisable statistical platform to infer the dynamic pathways by which many, potentially interacting, discrete traits are acquired or lost over time in biomedical systems. The platform uses HyperTraPS (hypercubic transition path sampling) to learn progression pathways from cross-sectional, longitudinal, or phylogenetically-linked data with unprecedented efficiency, readily distinguishing multiple competing pathways, and identifying the most parsimonious mechanisms underlying given observations. Its Bayesian structure quantifies uncertainty in pathway structure and allows interpretable predictions of behaviours, such as which symptom a patient will acquire next. We exploit the model's topology to provide visualisation tools for intuitive assessment of multiple, variable pathways. We apply the method to ovarian cancer progression and the evolution of multidrug resistance in tuberculosis, demonstrating its power to reveal previously undetected dynamic pathways.

q-bio.QM↗

Mitochondrial network state scales mtDNA genetic dynamics

Mitochondrial DNA (mtDNA) mutations cause severe congenital diseases but may also be associated with healthy aging. MtDNA is stochastically replicated and degraded, and exists within organelles which undergo dynamic fusion and fission. The role of the resulting mitochondrial networks in the time evolution of the cellular proportion of mutated mtDNA molecules (heteroplasmy), and cell-to-cell variability in heteroplasmy (heteroplasmy variance), remains incompletely understood. Heteroplasmy variance is particularly important since it modulates the number of pathological cells in a tissue. Here, we provide the first wide-reaching theoretical framework which bridges mitochondrial network and genetic states. We show that, under a range of conditions, the (genetic) rate of increase in heteroplasmy variance and de novo mutation are proportionally modulated by the (physical) fraction of unfused mitochondria, independently of the absolute fission-fusion rate. In the context of selective fusion, we show that intermediate fusion/fission ratios are optimal for the clearance of mtDNA mutants. Our findings imply that modulating network state, mitophagy rate and copy number to slow down heteroplasmy dynamics when mean heteroplasmy is low could have therapeutic advantages for mitochondrial disease and healthy aging.

q-bio.SC↗

Mitochondrial heterogeneity

Cell-to-cell heterogeneity drives a range of (patho)physiologically important phenomena, such as cell fate and chemotherapeutic resistance. The role of metabolism, and particularly mitochondria, is increasingly being recognised as an important explanatory factor in cell-to-cell heterogeneity. Most eukaryotic cells possess a population of mitochondria, in the sense that mitochondrial DNA (mtDNA) is held in multiple copies per cell, where the sequence of each molecule can vary. Hence intra-cellular mitochondrial heterogeneity is possible, which can induce inter-cellular mitochondrial heterogeneity, and may drive aspects of cellular noise. In this review, we discuss sources of mitochondrial heterogeneity (variations between mitochondria in the same cell, and mitochondrial variations between supposedly identical cells) from both genetic and non-genetic perspectives, and mitochondrial genotype-phenotype links. We discuss the apparent homeostasis of mtDNA copy number, the observation of pervasive intra-cellular mtDNA mutation (we term `microheteroplasmy') and developments in the understanding of inter-cellular mtDNA mutation (`macroheteroplasmy'). We point to the relationship between mitochondrial supercomplexes, cristal structure, pH and cardiolipin as a potential amplifier of the mitochondrial genotype-phenotype link. We also discuss mitochondrial membrane potential and networks as sources of mitochondrial heterogeneity, and their influence upon the mitochondrial genome. Finally, we revisit the idea of mitochondrial complementation as a means of dampening mitochondrial genotype-phenotype links in light of recent experimental developments. The diverse sources of mitochondrial heterogeneity, as well as their increasingly recognised role in contributing to cellular heterogeneity, highlights the need for future single-cell mitochondrial measurements in the context of cellular noise studies.

q-bio.SC↗

Endless love: On the termination of a playground number game

A simple and popular childhood game, `LOVES' or the `Love Calculator', involves an iterated rule applied to a string of digits and gives rise to surprisingly rich behaviour. Traditionally, players' names are used to set the initial conditions for an instance of the game: its behaviour for an exhaustive set of pairings of popular UK childrens' names, and for more general initial conditions, is examined. Convergence to a fixed outcome (the desired result) is not guaranteed, even for some plausible first name pairings. No pairs of top-50 common first names exhibit non-convergence, suggesting that it is rare in the playground; however, including surnames makes non-convergence more likely due to higher letter counts (for example, `Reese Witherspoon LOVES Calvin Harris'). Different game keywords (including from different languages) are also considered. An estimate for non-convergence propensity is derived: if the sum $m$ of digits in a string of length $w$ obeys $m > 18/(3/2)^{w-4}$, convergence is less likely. Pairs of top UK names with pairs of `O's and several `L's (for example, Chloe and Joseph, or Brooke and Scarlett) often attain high scores. When considering individual names playing with a range of partners, those with no `LOVES' letters score lowest, and names with intermediate (not simply the highest) letter counts often perform best, with Connor and Evie averaging the highest scores when played with other UK top names.

math.HO↗

Stochastic modelling, Bayesian inference, and new in vivo measurements elucidate the debated mtDNA bottleneck mechanism

Dangerous damage to mitochondrial DNA (mtDNA) can be ameliorated during mammalian development through a highly debated mechanism called the mtDNA bottleneck. Uncertainty surrounding this process limits our ability to address inherited mtDNA diseases. We produce a new, physically motivated, generalisable theoretical model for mtDNA populations during development, allowing the first statistical comparison of proposed bottleneck mechanisms. Using approximate Bayesian computation and mouse data, we find most statistical support for a combination of binomial partitioning of mtDNAs at cell divisions and random mtDNA turnover, meaning that the debated exact magnitude of mtDNA copy number depletion is flexible. New experimental measurements from a wild-derived mtDNA pairing in mice confirm the theoretical predictions of this model. We analytically solve a mathematical description of this mechanism, computing probabilities of mtDNA disease onset, efficacy of clinical sampling strategies, and effects of potential dynamic interventions, thus developing a quantitative and experimentally-supported stochastic theory of the bottleneck.

q-bio.QM↗

Efficient parametric inference for stochastic biological systems with measured variability

Stochastic systems in biology often exhibit substantial variability within and between cells. This variability, as well as having dramatic functional consequences, provides information about the underlying details of the system's behaviour. It is often desirable to infer properties of the parameters governing such systems given experimental observations of the mean and variance of observed quantities. In some circumstances, analytic forms for the likelihood of these observations allow very efficient inference: we present these forms and demonstrate their usage. When likelihood functions are unavailable or difficult to calculate, we show that an implementation of approximate Bayesian computation (ABC) is a powerful tool for parametric inference in these systems. However, the calculations required to apply ABC to these systems can also be computationally expensive, relying on repeated stochastic simulations. We propose an ABC approach that cheaply eliminates unimportant regions of parameter space, by addressing computationally simple mean behaviour before explicitly simulating the more computationally demanding variance behaviour. We show that this approach leads to a substantial increase in speed when applied to synthetic and experimental datasets.

q-bio.QM↗

Closed-form stochastic solutions for non-equilibrium dynamics and inheritance of cellular components over many cell divisions

Stochastic dynamics govern many important processes in cellular biology, and an underlying theoretical approach describing these dynamics is desirable to address a wealth of questions in biology and medicine. Mathematical tools exist for treating several important examples of these stochastic processes, most notably gene expression, and random partitioning at single cell divisions or after a steady state has been reached. Comparatively little work exists exploring different and specific ways that repeated cell divisions can lead to stochastic inheritance of unequilibrated cellular populations. Here we introduce a mathematical formalism to describe cellular agents that are subject to random creation, replication, and/or degradation, and are inherited according to a range of random dynamics at cell divisions. We obtain closed-form generating functions describing systems at any time after any number of cell divisions for binomial partitioning and divisions provoking a deterministic or random, subtractive or additive change in copy number, and show that these solutions agree exactly with stochastic simulation. We apply this general formalism to several example problems involving the dynamics of mitochondrial DNA (mtDNA) during development and organismal lifetimes.

q-bio.QM↗