Searcharxiv⌕ Search

arXiv subjects

Michael P. H. Stumpf

Publications and source records attributed to Michael P. H. Stumpf.

At least 19 recordsLinked to original sources

Phase transitions in microbial lineage trees

Microbial populations exhibit high cell-to-cell variability, which fundamentally shapes population behavior. A striking consequence is the existence of phase transitions, where small genetic or environmental changes trigger abrupt shifts in population dynamics. While biological phase transitions have often been proposed, connecting observed behavior to the underlying physics has remained challenging. We combine population genetics with statistical physics to show how phase transitions arise naturally in microbial populations. We highlight the existence of a first-order transition in a model of bacterial plasmid engineering and find a strict lower bound on the number of plasmids that can be stably maintained in a population.

q-bio.PE↗

Abstract relational structures in models of biology

The mathematical formalisms used to model biological systems induce both latent and ambiguous assumptions that can limit or distort their representational capabilities. Developing formalisms that can represent systems more precisely is fundamental to comprehending their intricacies and complexities. Here we introduce the systems hypergraph, a general and extendable formalism for representing abstract relational systems. A systems hypergraph combines a hypergraph, representing multidimensional relations among objects, with a hierarchical system of attributes representing system properties and their interdependencies. The attribute structure ensures that dependencies between system properties are patent and unambiguous, thereby clarifying assumptions and avoiding redundancy in data association. As an application we consider two formalisms widely used in systems biology - chemical reaction networks and stochastic Petri nets - and study their natural representation as systems hypergraphs. This allows us to relate the two formalisms rigorously, demonstrating in particular that stochastic Petri nets are strictly more general than chemical reaction networks in contrast to their commonly assumed equivalence. More broadly our work demonstrates the power of abstraction, and in particular its role in mediating between objects and relations in mathematical representations of biological complexity.

q-bio.QM↗

The two-clock problem in population dynamics

Biological time can be measured in two ways: in generations and in physical (chronological) time. When generations overlap, these two notions diverge, which impedes our ability to relate mathematical models to real populations. In this paper we show that nevertheless, the two clocks can be synchronised in the long run via a simple identity relating generational and physical time. This equivalence allows us to directly translate statements from the generational picture to the physical picture and vice versa. We derive a generalized Euler-Lotka equation linking the basic reproduction number $R_0$ to the growth rate, and present a simple identity that relates the selection coefficient of a mutation to the history of typical individuals, with applications to epidemiology, population biology and microbial growth.

q-bio.PE↗

FlowClass.jl: Classifying Dynamical Systems by Structural Properties in Julia

FlowClass.jl is a Julia package for classifying continuous-time dynamical systems into a hierarchy of structural classes: Gradient, Gradient-like, Morse-Smale, Structurally Stable, and General. Given a vector field \(\mathbf{F}(\mathbf{x})\) defining the system \(\mathrm{d}\mathbf{x}/\mathrm{d}t = \mathbf{F}(\mathbf{x})\), the package performs a battery of computational tests -- Jacobian symmetry analysis, curl magnitude estimation, fixed point detection and stability classification, periodic orbit detection, and stable/unstable manifold computation -- to determine where the system sits within the classification hierarchy. This classification has direct implications for qualitative behaviour: gradient systems cannot oscillate, Morse-Smale systems are structurally stable in less than 3 dimensions, and general systems may exhibit chaos. Much of classical developmental theory going back to Waddington's epigenetic landscape rests on an implicit assumption of gradient dynamics. The package is designed with applications in systems and developmental biology in mind, particularly the analysis of gene regulatory networks and cell fate decision models in the context of Waddington's epigenetic landscape. It provides tools to assess whether a landscape metaphor is appropriate for a given dynamical model, and to quantify the magnitude of non-gradient (curl) dynamics.

math.DS↗

Cell size distributions in lineages

Cells actively regulate their size during the cell cycle to maintain volume homeostasis across generations. While various mathematical models of cell size regulation have been proposed to explain how this is achieved, relating these models to experimentally observed cell size distributions has proved challenging. In this paper we present a simple formula for the cell size distribution in lineages as observed in e.g. a mother machine, and provide a new derivation for the corresponding result in populations, assuming exponential cell growth. Our results are independent of the underlying cell size control mechanism and explain the characteristic shape underlying experimentally observed cell size distributions. We furthermore derive universal moment identities for these distributions, and show that our predictions agree well with experimental measurements of E. coli cells, both on the distribution and the moment level.

q-bio.QM↗

Towards a mathematical framework for modelling cell fate dynamics

An adult human body is made up of some 30 to 40 trillion cells, all of which stem from a single fertilized egg cell. The process by which the right cells appear to arrive in their right numbers at the right time at the right place -- development -- is only understood in the roughest of outlines. This process does not happen in isolation: the egg, the embryo, the developing foetus, and the adult organism all interact intricately with their changing environments. Conceptual and, increasingly, mathematical approaches to modelling development have centred around Waddington's concept of an epigenetic landscape. This perspective enables us to talk about the molecular and cellular factors that contribute to cells reaching their terminally differentiated state: their fate. The landscape metaphor is however only a simplification of the complex process of development; it for instance does not consider environmental influences, a context which we argue needs to be explicitly taken into account and from the outset. When delving into the literature, it also quickly becomes clear that there is a lack of consistency and agreement on even fundamental concepts; for example, the precise meaning of what we refer to when talking about a `cell type' or `cell state.' Here we engage with previous theoretical and mathematical approaches to modelling cell fate -- focused on trees, networks, and landscape descriptions -- and argue that they require a level of simplification that can be problematic. We introduce random dynamical systems as one natural alternative. These provide a flexible conceptual and mathematical framework that is free of extraneous assumptions. We develop some of the basic concepts and discuss them in relation to now `classical' depictions of cell fate dynamics, in particular Waddington's landscape.

q-bio.SC↗

Mapping, modeling, and reprogramming cell-fate decision making systems

Many cellular processes involve information processing and decision making. We can probe these processes at increasing molecular detail. The analysis of heterogeneous data remains a challenge that requires new ways of thinking about cells in quantitative, predictive, and mechanistic ways. We discuss the role of mathematical models in the context of cell-fate decision making systems across the tree of life. Complex multi-cellular organisms have been a particular focus, but single celled organisms also have to sense and respond to their environment. We center our discussion around the idea of design principles which we can learn from observations and modeling, and exploit in order to (re)-design or guide cellular behavior.

q-bio.CB↗

Hypergraph Animals

Here we introduce simple structures for the analysis of complex hypergraphs, hypergraph animals. These structures are designed to describe the local node neighbourhoods of nodes in hypergraphs. We establish their relationships to lattice animals and network motifs, and we develop their combinatorial properties for sparse and uncorrelated hypergraphs. We make use of the tight link of hypergraph animals to partition numbers, which opens up a vast mathematical framework for the analysis of hypergraph animals. We then study their abundances in random hypergraphs. Two transferable insights result from this analysis: (i) it establishes the importance of high-cardinality edges in ensembles of random hypergraphs that are inspired by the classical Erdös-Renyí random graphs; and (ii) there is a close connection between degree and hyperedge cardinality in random hypergraphs that shapes animal abundances and spectra profoundly. Both findings imply that hypergraph animals can have the potential to affect information flow and processing in complex systems. Our analysis of also suggests that we need to spend more effort on investigating and developing suitable conditional ensembles of random hypergraphs that can capture real-world structures and their complex dependency structures.

q-bio.MN↗

Hypergraphs for multiscale cycles in structured data

Scientific data has been growing in both size and complexity across the modern physical, engineering, life and social sciences. Spatial structure, for example, is a hallmark of many of the most important real-world complex systems, but its analysis is fraught with statistical challenges. Topological data analysis can provide a powerful computational window on complex systems. Here we present a framework to extend and interpret persistent homology summaries to analyse spatial data across multiple scales. We introduce hyperTDA, a topological pipeline that unifies local (e.g. geodesic) and global (e.g. Euclidean) metrics without losing spatial information, even in the presence of noise. Homology generators offer an elegant and flexible description of spatial structures and can capture the information computed by persistent homology in an interpretable way. Here the information computed by persistent homology is transformed into a weighted hypergraph, where hyperedges correspond to homology generators. We consider different choices of generators (e.g. matroid or minimal) and find that centrality and community detection are robust to either choice. We compare hyperTDA to existing geometric measures and validate its robustness to noise. We demonstrate the power of computing higher-order topological structures on spatial curves arising frequently in ecology, biophysics, and biology, but also in high-dimensional financial datasets. We find that hyperTDA can select between synthetic trajectories from the landmark 2020 AnDi challenge and quantifies movements of different animal species, even when data is limited.

math.AT↗

A group theoretic approach to model comparison with simplicial representations

The complexity of biological systems, and the increasingly large amount of associated experimental data, necessitates that we develop mathematical models to further our understanding of these systems. As biological systems are generally not well understood, most mathematical models of these systems are based on experimental data, resulting in a seemingly heterogeneous collection of models that ostensibly represent the same system. To understand the system we therefore need to know how the different models are related, with a view to obtaining a unified mathematical description. This goal is complicated by the fact that distinct mathematical formalisms may be used to represent the same system, making direct comparison of the models very difficult. In previous work we developed an appropriate framework for model comparison where we represent models as labelled simplicial complexes and compare them with two general methodologies: comparison by distance or equivalence. In this article we continue the development of our model comparison methodology in two directions. First, we present a rigorous and automatable methodology for the core process of comparison by equivalence, namely determining the vertices in a simplicial representation, corresponding to model components, that are conceptually related and the identification of these vertices via simplicial operations. Our methodology is based on considerations of vertex symmetry in the simplicial representation, for which we develop the required mathematical theory of group actions on simplicial complexes. This methodology greatly simplifies and expedites the process of determining model equivalence. Second, we provide an alternative mathematical framework for our model-comparison methodology by representing models as groups, which allows for the direct application of group-theoretic techniques within our model-comparison methodology.

q-bio.QM↗

Open Problems in Mathematical Biology

Biology is data-rich, and it is equally rich in concepts and hypotheses. Part of trying to understand biological processes and systems is therefore to confront our ideas and hypotheses with data using statistical methods to determine the extent to which our hypotheses agree with reality. But doing so in a systematic way is becoming increasingly challenging as our hypotheses become more detailed, and our data becomes more complex. Mathematical methods are therefore gaining in importance across the life- and biomedical sciences. Mathematical models allow us to test our understanding, make testable predictions about future behaviour, and gain insights into how we can control the behaviour of biological systems. It has been argued that mathematical methods can be of great benefit to biologists to make sense of data. But mathematics and mathematicians are set to benefit equally from considering the often bewildering complexity inherent to living systems. Here we present a small selection of open problems and challenges in mathematical biology. We have chosen these open problems because they are of both biological and mathematical interest.

q-bio.QM↗

Julia for Biologists

Increasing emphasis on data and quantitative methods in the biomedical sciences is making biological research more computational. Collecting, curating, processing, and analysing large genomic and imaging data sets poses major computational challenges, as does simulating larger and more realistic models in systems biology. Here we discuss how a relative newcomer among computer programming languages -- Julia -- is poised to meet the current and emerging demands in the computational biosciences, and beyond. Speed, flexibility, a thriving package ecosystem, and readability are major factors that make high-performance computing and data analysis available to an unprecedented degree to "gifted amateurs". We highlight how Julia's design is already enabling new ways of analysing biological data and systems, and we provide a, necessarily incomplete, list of resources that can facilitate the transition into the Julian way of computing.

q-bio.QM↗

Model comparison via simplicial complexes and persistent homology

In many scientific and technological contexts we have only a poor understanding of the structure and details of appropriate mathematical models. We often, therefore, need to compare different models. With available data we can use formal statistical model selection to compare and contrast the ability of different mathematical models to describe such data. There is, however, a lack of rigorous methods to compare different models \emph{a priori}. Here we develop and illustrate two such approaches that allow us to compare model structures in a systematic way {by representing models in terms of simplicial complexes}. Using well-developed concepts from simplicial algebraic topology, we define a distance between models based on their simplicial representations. Employing persistent homology with a flat filtration provides for alternative representations of the models as persistence intervals, which represent the structure of the models, from which we can also obtain the distances between models. We then expand on this measure of model distance to study the concept of model equivalence in order to determine the conceptual similarity of models. We apply our methodology for model comparison to demonstrate an equivalence between a positional-information model and a Turing-pattern model from developmental biology, constituting a novel observation for two classes of models that were previously regarded as unrelated.

math.AT↗

Bayesian and Algebraic Strategies to Design in Synthetic Biology

Innovation in synthetic biology often still depends on large-scale experimental trial-and-error, domain expertise, and ingenuity. The application of rational design engineering methods promise to make this more efficient, faster, cheaper and safer. But this requires mathematical models of cellular systems. And for these models we then have to determine if they can meet our intended target behaviour. Here we develop two complementary approaches that allow us to determine whether a given molecular circuit, represented by a mathematical model, is capable of fulfilling our design objectives. We discuss algebraic methods that are capable of identifying general principles guaranteeing desired behaviour; and we provide an overview over Bayesian design approaches that allow us to choose from a set of models, that model which has the highest probability of fulfilling our design objectives. We discuss their uses in the context of biochemical adaptation, and then consider how robustness can and should affect our design approach.

q-bio.MN↗

Non-equilibrium statistical physics, transitory epigenetic landscapes, and cell fate decision dynamics

Statistical physics provides a useful perspective for the analysis of many complex systems; it allows us to relate microscopic fluctuations to macroscopic observations. Developmental biology, but also cell biology more generally, are examples where apparently robust behaviour emerges from highly complex and stochastic sub-cellular processes. Here we attempt to make connections between different theoretical perspectives to gain qualitative insights into the types of cell-fate decision making processes that are at the heart of stem cell and developmental biology. We discuss both dynamical systems as well as statistical mechanics perspectives on the classical Waddington or epigenetic landscape. We find that non-equilibrium approaches are required to overcome some of the shortcomings of classical equilibrium statistical thermodynamics or statistical mechanics in order to shed light on biological processes, which, almost by definition, are typically far from equilibrium.

q-bio.CB↗

Mathematical and Statistical Techniques for Systems Medicine: The Wnt Signaling Pathway as a Case Study

The last decade has seen an explosion in models that describe phenomena in systems medicine. Such models are especially useful for studying signaling pathways, such as the Wnt pathway. In this chapter we use the Wnt pathway to showcase current mathematical and statistical techniques that enable modelers to gain insight into (models of) gene regulation, and generate testable predictions. We introduce a range of modeling frameworks, but focus on ordinary differential equation (ODE) models since they remain the most widely used approach in systems biology and medicine and continue to offer great potential. We present methods for the analysis of a single model, comprising applications of standard dynamical systems approaches such as nondimensionalization, steady state, asymptotic and sensitivity analysis, and more recent statistical and algebraic approaches to compare models with data. We present parameter estimation and model comparison techniques, focusing on Bayesian analysis and coplanarity via algebraic geometry. Our intention is that this (non exhaustive) review may serve as a useful starting point for the analysis of models in systems medicine.

q-bio.QM↗

Conditional Random Matrix Ensembles and the Stability of Dynamical Systems

There has been a long-standing and at times fractious debate whether complex and large systems can be stable. In ecology, the so-called `diversity-stability debate' arose because mathematical analyses of ecosystem stability were either specific to a particular model (leading to results that were not general), or chosen for mathematical convenience, yielding results unlikely to be meaningful for any interesting realistic system. May's work, and its subsequent elaborations, relied upon results from random matrix theory, particularly the circular law and its extensions, which only apply when the strengths of interactions between entities in the system are assumed to be independent and identically distributed (i.i.d.). Other studies have optimistically generalised from the analysis of very specific systems, in a way that does not hold up to closer scrutiny. We show here that this debate can be put to rest, once these two contrasting views have been reconciled --- which is possible in the statistical framework developed here. Here we use a range of illustrative examples of dynamical systems to demonstrate that (i) stability probability cannot be summarily deduced from any single property of the system (e.g. its diversity), and (ii) our assessment of stability depends on adequately capturing the details of the systems analysed. Failing to condition on the structure of dynamical systems will skew our analysis and can, even for very small systems, result in an unnecessarily pessimistic diagnosis of their stability.

math.DS↗

StochDecomp - Matlab package for noise decomposition in stochastic biochemical systems

Stochasticity is an indispensable aspect of biochemical processes at the cellular level. Studies on how the noise enters and propagates in biochemical systems provided us with nontrivial insights into the origins of stochasticity, in total however they constitute a patchwork of different theoretical analyses. Here we present a flexible and generally applicable noise decomposition tool, that allows us to calculate contributions of individual reactions to the total variability of a system's output. With the package it is therefore possible to quantify how the noise enters and propagates in biochemical systems. We also demonstrate and exemplify using the JAK-STAT signalling pathway that it is possible to infer noise contributions resulting from individual reactions directly from experimental data. This is the first computational tool that allows to decompose noise into contributions resulting from individual reactions.

q-bio.QM↗