Searcharxiv⌕ Search

arXiv subjects

Mike Steel

Publications and source records attributed to Mike Steel.

At least 37 records · Page 2Linked to original sources

Combinatorics of polymer models of early metabolism

Polymer models are a widely used tool to study the prebiotic formation of metabolism at the origins of life. Counts of the number of reactions in these models are often crucial in probabilistic arguments concerning the emergence of autocatalytic networks. In the first part of this paper, we provide the first exact description of the number of reactions under widely applied model assumptions. Conclusions from earlier studies rely on either approximations or asymptotic counting, and we show that the exact counts lead to similar, though not always identical, asymptotic results. In the second part of the paper, we investigate a novel model assumption whereby polymers are invariant under spatial rotation. We outline the biochemical relevance of this condition and again give exact enumerative and asymptotic formulae for the number of reactions.

q-bio.MN↗

Defining phylogenetic networks using ancestral profiles

Rooted phylogenetic networks provide a more complete representation of the ancestral relationship between species than phylogenetic trees when reticulate evolutionary processes are at play. One way to reconstruct a phylogenetic network is to consider its `ancestral profile' (the number of paths from each ancestral vertex to each leaf). In general, this information does not uniquely determine the underlying phylogenetic network. A recent paper considered a new class of phylogenetic networks called `orchard networks' where this uniqueness was claimed to hold. Here we show that an additional restriction on the network, that of being `stack-free', is required in order for the original uniqueness claim to hold. On the other hand, if the additional stack-free restriction is lifted, we establish an alternative result; namely, there is uniqueness within the class of orchard networks up to the resolution of vertices of high in-degree.

math.CO↗

The expected number of viable autocatalytic sets in chemical reaction systems

The emergence of self-sustaining autocatalytic networks in chemical reaction systems has been studied as a possible mechanism for modelling how living systems first arose. It has been known for several decades that such networks will form within systems of polymers (under cleavage and ligation reactions) under a simple process of random catalysis, and this process has since been mathematically analysed. In this paper, we provide an exact expression for the expected number of self-sustaining autocatalytic networks that will form in a general chemical reaction system, and the expected number of these networks that will also be uninhibited (by some molecule produced by the system). Using these equations, we are able to describe the patterns of catalysis and inhibition that maximise or minimise the expected number of such networks. We apply our results to derive a general theorem concerning the trade-off between catalysis and inhibition, and to provide some insight into the extent to which the expected number of self-sustaining autocatalytic networks coincides with the probability that at least one such system is present.

q-bio.MN↗

Modeling a Cognitive Transition at the Origin of Cultural Evolution using Autocatalytic Networks

Autocatalytic networks have been used to model the emergence of self-organizing structure capable of sustaining life and undergoing biological evolution. Here, we model the emergence of cognitive structure capable of undergoing cultural evolution. Mental representations of knowledge and experiences play the role of catalytic molecules, and interactions amongst them (e.g., the forging of new associations) play the role of reactions, and result in representational redescription. The approach tags mental representations with their source, i.e., whether they were acquired through social learning, individual learning (of pre-existing information), or creative thought (resulting in the generation of new information). This makes it possible to model how cognitive structure emerges, and to trace lineages of cumulative culture step by step. We develop a formal representation of the cultural transition from Oldowan to Acheulean tool technology using Reflexively Autocatalytifc and Food set generated (RAF) networks. Unlike more primitive Oldowan stone tools, the Acheulean hand axe required not only the capacity to envision and bring into being something that did not yet exist, but hierarchically structured thought and action, and the generation of new mental representations: the concepts EDGING, THINNING, SHAPING, and a meta-concept, HAND AXE. We show how this constituted a key transition towards the emergence of semantic networks that were self-organizing, self-sustaining, and autocatalytic, and discuss how such networks replicated through social interaction. The model provides a promising approach to unraveling one of the greatest anthropological mysteries: that of why development of the Acheulean hand axe was followed by over a million years of cultural stasis.

physics.soc-ph↗

Combinatorial results for network-based models of metabolic origins

A key step in the origin of life is the emergence of a primitive metabolism. This requires the formation of a subset of chemical reactions that is both self-sustaining and collectively autocatalytic. A generic theory to study such processes (called 'RAF theory') has provided a precise and computationally effective way to address these questions, both on simulated data and in laboratory studies. One of the classic applications of this theory (arising from Stuart Kauffman's pioneering work in the 1980s) involves networks of polymers under cleavage and ligation reactions; in the first part of this paper, we provide the first exact description of the number of such reactions under various model assumptions. Conclusions from earlier studies relied on either approximations or asymptotic counting, and we show that the exact counts lead to similar (though not always identical) asymptotic results. In the second part of the paper, we solve some questions posed in more recent papers concerning the computational complexity of some key questions in RAF theory. In particular, although there is a fast algorithm to determine whether or not a catalytic reaction network contains a subset that is both self-sustaining and autocatalytic (and, if so, find one), determining whether or not sets exist that satisfy certain additional constraints exist turns out to be NP-complete.

q-bio.MN↗

Combinatorial properties of phylogenetic diversity indices

Phylogenetic diversity indices provide a formal way to apportion 'evolutionary heritage' across species. Two natural diversity indices are Fair Proportion (FP) and Equal Splits (ES). FP is also called 'evolutionary distinctiveness' and, for rooted trees, is identical to the Shapley Value (SV), which arises from cooperative game theory. In this paper, we investigate the extent to which FP and ES can differ, characterise tree shapes on which the indices are identical, and study the equivalence of FP and SV and its implications in more detail. We also define and investigate analogues of these indices on unrooted trees (where SV was originally defined), including an index that is closely related to the Pauplin representation of phylogenetic diversity.

q-bio.PE↗

Dynamics of a birth-death process based on combinatorial innovation

A feature of human creativity is the ability to take a subset of existing items (e.g. objects, ideas, or techniques) and combine them in various ways to give rise to new items, which, in turn, fuel further growth. Occasionally, some of these items may also disappear (extinction). We model this process by a simple stochastic birth--death model, with non-linear combinatorial terms in the growth coefficients to capture the propensity of subsets of items to give rise to new items. In its simplest form, this model involves just two parameters $(P, α)$. This process exhibits a characteristic 'hockey-stick' behaviour: a long period of relatively little growth followed by a relatively sudden 'explosive' increase. We provide exact expressions for the mean and variance of this time to explosion and compare the results with simulations. We then generalise our results to allow for more general parameter assignments, and consider possible applications to data involving human productivity and creativity.

q-bio.PE↗

A class of phylogenetic networks reconstructable from ancestral profiles

Rooted phylogenetic networks provide an explicit representation of the evolutionary history of a set $X$ of sampled species. In contrast to phylogenetic trees which show only speciation events, networks can also accommodate reticulate processes (for example, hybrid evolution, endosymbiosis, and lateral gene transfer). A major goal in systematic biology is to infer evolutionary relationships, and while phylogenetic trees can be uniquely determined from various simple combinatorial data on $X$, for networks the reconstruction question is much more subtle. Here we ask when can a network be uniquely reconstructed from its `ancestral profile' (the number of paths from each ancestral vertex to each element in $X$). We show that reconstruction holds (even within the class of all networks) for a class of networks we call `orchard networks', and we provide a polynomial-time algorithm for reconstructing any orchard network from its ancestral profile. Our approach relies on establishing a structural theorem for orchard networks, which also provides for a fast (polynomial-time) algorithm to test if any given network is of orchard type. Since the class of orchard networks includes tree-sibling tree-consistent networks and tree-child networks, our result generalise reconstruction results from 2008 and 2009. Orchard networks allow for an unbounded number $k$ of reticulation vertices, in contrast to tree-sibling tree-consistent networks and tree-child networks for which $k$ is at most $2|X|-4$ and $|X|-1$, respectively.

math.CO↗

Tree-based networks: characterisations, metrics, and support trees

Phylogenetic networks generalise phylogenetic trees and allow for the accurate representation of the evolutionary history of a set of present-day species whose past includes reticulate events such as hybridisation and lateral gene transfer. One way to obtain such a network is by starting with a (rooted) phylogenetic tree $T$, called a base tree, and adding arcs between arcs of $T$. The class of phylogenetic networks that can be obtained in this way is called tree-based networks and includes the prominent classes of tree-child and reticulation-visible networks. Initially defined for binary phylogenetic networks, tree-based networks naturally extend to arbitrary phylogenetic networks. In this paper, we generalise recent tree-based characterisations and associated proximity measures for binary phylogenetic networks to arbitrary phylogenetic networks. These characterisations are in terms of matchings in bipartite graphs, path partitions, and antichains. Some of the generalisations are straightforward to establish using the original approach, while others require a very different approach. Furthermore, for an arbitrary tree-based network $N$, we characterise the support trees of $N$, that is, the tree-based embeddings of $N$. We use this characterisation to give an explicit formula for the number of support trees of $N$ when $N$ is binary. This formula is written in terms of the components of a bipartite graph.

q-bio.PE↗

Quantifying the accuracy of ancestral state prediction in a phylogenetic tree under maximum parsimony

In phylogenetic studies, biologists often wish to estimate the ancestral discrete character state at an interior vertex $v$ of an evolutionary tree $T$ from the states that are observed at the leaves of the tree. A simple and fast estimation method --- maximum parsimony --- takes the ancestral state at $v$ to be any state that minimises the number of state changes in $T$ required to explain its evolution on $T$. In this paper, we investigate the reconstruction accuracy of this estimation method further, under a simple symmetric model of state change, and obtain a number of new results, both for 2-state characters, and $r$--state characters ($r>2$). Our results rely on establishing new identities and inequalities, based on a coupling argument that involves a simpler `coin toss' approach to ancestral state reconstruction.

q-bio.PE↗

Tractable models of self-sustaining autocatalytic networks

Self-sustaining autocatalytic networks play a central role in living systems, from metabolism at the origin of life, simple RNA networks, and the modern cell, to ecology and cognition. A collectively autocatalytic network that can be sustained from an ambient food set is also referred to more formally as a `Reflexively Autocatalytic F-generated' (RAF) set. In this paper, we first investigate a simplified setting for studying RAFs, which are nevertheless relevant to real biochemistry and allows for a more exact mathematical analysis based on graph-theoretic concepts. This, in turn, allows for the development of efficient (polynomial-time) algorithms for questions that are computationally NP-hard in the general RAF setting. We then show how this simplified setting for RAF systems leads naturally to a more general notion of RAFs that are `generative' (they can be built up from simpler RAFs) and for which efficient algorithms carry over to this more general setting. Finally, we show how classical RAF theory can be extended to deal with ensembles of catalysts as well as the assignment of rates to reactions according to which catalysts (or combinations of catalysts) are available.

q-bio.MN↗

Phylogenetic flexibility via Hall-type inequalities and submodularity

Given a collection $τ$ of subsets of a finite set $X$, we say that $τ$ is {\em phylogenetically flexible} if, for any collection $R$ of rooted phylogenetic trees whose leaf sets comprise the collection $τ$, $R$ is compatible (i.e. there is a rooted phylogenetic $X$--tree that displays each tree in $R$). We show that $τ$ is phylogenetically flexible if and only if it satisfies a Hall-type inequality condition of being `slim'. Using submodularity arguments, we show that there is a polynomial-time algorithm for determining whether or not $τ$ is slim. This `slim' condition reduces to a simpler inequality in the case where all of the sets in $τ$ have size 3, a property we call `thin'. Thin sets were recently shown to be equivalent to the existence of an (unrooted) tree for which the median function provides an injective mapping to its vertex set; we show here that the unrooted tree in this representation can always be chosen to be a caterpillar tree. We also characterise when a collection $τ$ of subsets of size 2 is thin (in terms of the flexibility of total orders rather than phylogenies) and show that this holds if and only if an associated bipartite graph is a forest. The significance of our results for phylogenetics is in providing precise and efficiently verifiable conditions under which supertree methods that require consistent inputs of trees, can be applied to any input trees on given subsets of species.

math.CO↗

On the information content of discrete phylogenetic characters

Phylogenetic inference aims to reconstruct the evolutionary relationships of different species based on genetic (or other) data. Discrete characters are a particular type of data, which contain information on how the species should be grouped together. However, it has long been known that some characters contain more information than others. For instance, a character that assigns the same state to each species groups all of them together and so provides no insight into the relationships of the species considered. At the other extreme, a character that assigns a different state to each species also conveys no phylogenetic signal. In this manuscript, we study a natural combinatorial measure of the information content of an individual character and analyse properties of characters that provide the maximum phylogenetic information, particularly, the number of states such a character uses and how the different states have to be distributed among the species or taxa of the phylogenetic tree.

q-bio.PE↗

Species notions that combine phylogenetic trees and phenotypic partitions

A recent paper (Manceau and Lambert, 2016) developed a novel approach for describing two well-defined notions of 'species' based on a phylogenetic tree and a phenotypic partition. In this paper, we explore some further combinatorial properties of this approach and describe an extension that allows an arbitrary number of phenotypic partitions to be combined with a phylogenetic tree for these two species notions.

q-bio.PE↗

New Characterisations of Tree-Based Networks and Proximity Measures

Phylogenetic networks are a type of directed acyclic graph that represent how a set $X$ of present-day species are descended from a common ancestor by processes of speciation and reticulate evolution. In the absence of reticulate evolution, such networks are simply phylogenetic (evolutionary) trees. Moreover, phylogenetic networks that are not trees can sometimes be represented as phylogenetic trees with additional directed edges placed between their edges. Such networks are called {\em tree based}, and the class of phylogenetic networks that are tree based has recently been characterised. In this paper, we establish a number of new characterisations of tree-based networks in terms of path partitions and antichains (in the spirit of Dilworth's theorem), as well as via matchings in a bipartite graph. We also show that a temporal network is tree based if and only if it satisfies an antichain-to-leaf condition. In the second part of the paper, we define three indices that measure the extent to which an arbitrary phylogenetic network deviates from being tree based. We describe how these three indices can be described exactly and computed efficiently using classical results concerning maximum-sized matchings in bipartite graphs.

math.CO↗

Autocatalytic networks in cognition and the origin of culture

It has been proposed that cultural evolution was made possible by a cognitive transition brought about by onset of the capacity for self-triggered recall and rehearsal. Here we develop a novel idea that models of collectively autocatalytic networks, developed for understanding the origin and organization of life, may also help explain the origin of the kind of cognitive structure that makes cultural evolution possible. In our setting, mental representations (for example, memories, concepts, ideas) play the role of 'molecules', and 'reactions' involve the evoking of one representation by another through remindings, associations, and stimuli. In the 'episodic mind', representations are so coarse-grained (encode too few properties) that such reactions are catalyzed only by external stimuli. As cranial capacity increased, representations became more fine-grained (encoded more features), allowing them to act as catalysts, leading to streams of thought. At this point, the mind could combine representations and adapt them to specific needs and situations, and thereby contribute to cultural evolution. In this paper, we propose and study a simple and explicit cognitive model that gives rise naturally to autocatylatic networks, and thereby provides a possible mechanism for the transition from a pre-cultural episodic mind to a mimetic mind.

q-bio.NC↗

Combinatorial properties of triplet covers for binary trees

It is a classical result that an unrooted tree $T$ having positive real-valued edge lengths and no vertices of degree two can be reconstructed from the induced distance between each pair of leaves. Moreover, if each non-leaf vertex of $T$ has degree 3 then the number of distance values required is linear in the number of leaves. A canonical candidate for such a set of pairs of leaves in $T$ is the following: for each non-leaf vertex $v$, choose a leaf in each of the three components of $T-v$, group these three leaves into three pairs, and take the union of this set over all choices of $v$. This forms a so-called 'triplet cover' for $T$. In the first part of this paper we answer an open question (from 2012) by showing that the induced leaf-to-leaf distances for any triplet cover for $T$ uniquely determine $T$ and its edge lengths. We then investigate the finer combinatorial properties of triplet covers. In particular, we describe the structure of triplet covers that satisfy one or more of the following properties of being minimal, 'sparse', and 'shellable'.

math.CO↗

The optimal rate for resolving a near-polytomy in a phylogeny

The reconstruction of phylogenetic trees from discrete character data typically relies on models that assume the characters evolve under a continuous-time Markov process operating at some overall rate $λ$. When $λ$ is too high or too low, it becomes difficult to distinguish a short interior edge from a polytomy (the tree that results from collapsing the edge). In this note, we investigate the rate that maximizes the expected log-likelihood ratio (i.e. the Kullback--Leibler separation) between the four-leaf unresolved (star) tree and a four-leaf binary tree with interior edge length $ε$. For a simple two-state model, we show that as $ε$ converges to $0$ the optimal rate also converges to zero when the four pendant edges have equal length. However, when the four pendant branches have unequal length, two local optima can arise, and it is possible for the globally optimal rate to converge to a non-zero constant as $ε\rightarrow 0$. Moreover, in the setting where the four pendant branches have equal lengths and either (i) we replace the two-state model by an infinite-state model or (ii) we retain the two-state model and replace the Kullback--Leibler separation by Euclidean distance as the maximization goal, then the optimal rate also converges to a non-zero constant.

q-bio.PE↗