SearcharxivSearch

arXiv subjects

Mike Steel

Publications and source records attributed to Mike Steel.

At least 19 recordsLinked to original sources

Properties of biodiversity indices that incorporate future extinction risk

The loss of biodiversity due to the likely widespread extinction of species in the near future is a focus of current concern in conservation biology. One approach to measure the impact of this extinction is based on the predicted loss of phylogenetic diversity. These predictions have become a focus of the Zoological Society of London's `EDGE2' program for quantifying biodiversity loss and involves considering the HED (heightened evolutionary distinctiveness) and HEDGE (heightened evolutionary distinctiveness and globally endangered) indices which are based on phylogenetic diversity on a tree. Here, we show how to generalise the HED(GE) indices by expanding their application to more general settings (to phylogenetic networks, to feature diversity on discrete traits, and to arbitrary biodiversity measures). We provide a simple and explicit description of the mean and, importantly, the variance of such measures, and illustrate our results by an application to the phylogeny and a small set of features for all 27 extant Crocodilians.

q-bio.PE

Intermediate stages in the origin of metabolism at a phosphorylating hydrothermal vent

The origin of life required the emergence of metabolism, an autocatalytic network of enzymatic reactions that synthesize amino acids, nucleotides and cofactors. At the origin of metabolism there were no enzymes--how did it start? Empirical studies addressing early metabolic evolution are lacking. Harnessing protein structures for metabolic enzymes, we identify intermediate states in primordial metabolic assembly. We show that enzymatic metabolism in the universal common ancestor was incomplete, undergoing final assembly independently in the lineages leading to Bacteria and Archaea. Native transition metals--Fe0, Co0, Ni0, Pd0--served as the catalytic forerunners of both enzymes and cofactors at metabolic origin while phosphite supplied energy, as it phosphorylates AMP to ADP and serine to phosphoserine using native metal catalysts in water. Phosphite and native metals occur in serpentinizing hydrothermal systems, identifying an energy-supplying, catalytic site of metabolic origin. Cofactors liberated nascent metabolism from native metal catalysts, engendering its autocatalytic state.

q-bio.PE

A dichotomy law for certain classes of phylogenetic networks

Many classes of phylogenetic networks have been proposed in the literature. A feature of several of these classes is that if one restricts a network in the class to a subset of its leaves, then the resulting network may no longer lie within this class. This has implications for their biological applicability, since some species -- which are the leaves of an underlying evolutionary network -- may be missing (e.g., they may have become extinct, or there are no data available for them) or we may simply wish to focus attention on a subset of the species. On the other hand, certain classes of networks are `closed' when we restrict to subsets of leaves, such as (i) the classes of all phylogenetic networks or all phylogenetic trees; (ii) the classes of galled networks, simplicial networks, galled trees; and (iii) the classes of networks that have some parameter that is monotone-under-leaf-subsampling (e.g., the number of reticulations, height, etc.) bounded by some fixed value. It is easily shown that a closed subclass of phylogenetic trees is either all trees or a vanishingly small proportion of them (as the number of leaves grows). In this short paper, we explore whether this dichotomy phenomenon holds for other classes of phylogenetic networks, and their subclasses.

q-bio.PE

Counting rankings of tree-child networks

Rooted phylogenetic networks allow biologists to represent evolutionary relationships between present-day species by revealing ancestral speciation and hybridization events. A convenient and well-studied class of such networks are `tree-child networks' and a `ranking' of such a network is a temporal ordering of the ancestral speciation and hybridization events. In this short note, we investigate the question of counting such rankings on any given binary (or semi-binary) tree-child network. We also consider a class of binary tree-child networks that have exactly one ranking, and investigate further the relationship between ranked-tree child networks and the class of `normal' networks. Finally, we provide an explicit asymptotic expression for the expected number of rankings of a tree-child network chosen uniformly at random.

q-bio.PE

Predicting the depth of the most recent common ancestor of a random sample of $k$ species: the impact of phylogenetic tree shape

We consider the following question: how close to the ancestral root of a phylogenetic tree is the most recent common ancestor of $k$ species randomly sampled from the tips of the tree? For trees having shapes predicted by the Yule-Harding model, it is known that the most recent common ancestor is likely to be close to (or equal to) the root of the full tree, even as $n$ becomes large (for $k$ fixed). However, this result does not extend to models of tree shape that more closely describe phylogenies encountered in evolutionary biology. We investigate the impact of tree shape (via the Aldous $\beta-$splitting model) to predict the number of edges that separate the most recent common ancestor of a random sample of $k$ tip species and the root of the parent tree they are sampled from. Both exact and asymptotic results are presented. We also briefly consider a variation of the process in which a random number of tip species are sampled.

q-bio.PE

The asymptotic distribution of the $k$-Robinson-Foulds dissimilarity measure on labelled trees

Motivated by applications in medical bioinformatics, Khayatian et al. (2024) introduced a family of metrics on Cayley trees (the $k$-RF distance, for $k=0, \ldots, n-2$) and explored their distribution on pairs of random Cayley trees via simulations. In this paper, we investigate this distribution mathematically, and derive exact asymptotic descriptions of the distribution of the $k$-RF metric for the extreme values $k=0$ and $k=n-2$, as $n$ becomes large. We show that a linear transform of the $0$-RF metric converges to a Poisson distribution (with mean 2) whereas a similar transform for the $(n-2)$-RF metric leads to a normal distribution (with mean $\sim ne^{-2}$). These results (together with the case $k=1$ which behaves quite differently, and $k=n-3$) shed light on the earlier simulation results, and the predictions made concerning them.

math.PR

Asymptotic enumeration of normal and hybridization networks via tree decoration

Phylogenetic networks provide a more general description of evolutionary relationships than rooted phylogenetic trees. One way to produce a phylogenetic network is to randomly place $k$ arcs between the edges of a rooted binary phylogenetic tree with $n$ leaves. The resulting directed graph may fail to be a phylogenetic network, and even when it is (and thereby a `tree-based' network), it may fail to be a tree-child or normal network. In this paper, we first show that if $k$ is fixed, the proportion of arc placements that result in a normal network tends to 1 as $n$ grows. From this result, the asymptotic enumeration of normal networks becomes straightforward and provides a transparent meaning to the combinatorial terms that arise. Moreover, the approach extends to allow $k$ to grow with $n$ (at the rate $o(n^\frac{1}{3})$), which was not handled in earlier work. We also investigate a subclass of normal networks of particular relevance in biology (hybridization networks) and establish that the same asymptotic results apply.

q-bio.PE

Cumulative, Adaptive, Open-ended Change through Self-Other Reorganization: Reply to comment on 'An evolutionary process without variation and selection'

Self-Other Reorganization (SOR) is a theory of how interacting entities or individuals, each of which can be described as an autocatalytic network, collectively exhibit cumulative, adaptive, open-ended change, or evolution. Zachar et al.'s critique of SOR stems from misunderstandings; it does not weaken the arguments in (Gabora & Steel, 2021). The formal framework of Reflexively Autocatalytic and foodset-derived sets (RAFs) enables us to model the process whereby, through their interactions, a set of elements become a 'collective self.' SOR shows how the RAF setting provides a means of encompassing abiogenesis and cultural evolution under the same explanatory framework and provides a plausible explanation for the origins of both evolutionary processes. Although SOR allows for detrimental stimuli (and products), there is (naturally) limited opportunity for elements that do not contribute to or reinforce a RAF to become part of it. Replication and cumulative, adaptive change in RAFs is well-established in the literature. Contrary to Zachar et al., SOR is not a pure percolation model (such as SIR); it encompasses not only learning (modeled as assimilation of foodset elements) but also creative restructuring (modeled as generation of foodset-derived elements), as well as the emergence of new structures made possible by new foodset- and foodset-derived elements. Cultural SOR is robust to degradation, and imperfect replication. Zachar et al.'s simulation contains no RAFs, and does not model SOR.

q-bio.PE

An evolutionary process without variation and selection

Natural selection successfully explains how organisms accumulate adaptive change despite that traits acquired over a lifetime are eliminated at the end of each generation. However, in some domains that exhibit cumulative, adaptive change -- e.g., cultural evolution, and earliest life -- acquired traits are retained; these domains do not face the problem that Darwin's theory was designed to solve. Lack of transmission of acquired traits occurs when germ cells are protected from environmental change, due to a self-assembly code used in two distinct ways: (i) actively interpreted during development to generate a soma, and (ii) passively copied without interpretation during reproduction to generate germ cells. Early life and cultural evolution appear not to involve a self-assembly code used in these two ways. We suggest that cumulative, adaptive change in these domains is due to a lower-fidelity evolutionary process, and model it using Reflexively Autocatalytic and Foodset-generated networks. We refer to this more primitive evolutionary process as Self-Other Reorganisation (SOR) because it involves internal self-organising and self-maintaining processes within entities, as well as interaction between entities. SOR encompasses learning but in general operates across groups. We discuss the relationship between SOR and Lamarckism, and illustrate a special case of SOR without variation.

q-bio.PE

Transformations to simplify phylogenetic networks

The evolutionary relationships between species are typically represented in the biological literature by rooted phylogenetic trees. However, a tree fails to capture ancestral reticulate processes, such as the formation of hybrid species or lateral gene transfer events between lineages, and so the history of life is more accurately described by a rooted phylogenetic network. Nevertheless, phylogenetic networks may be complex and difficult to interpret, so biologists sometimes prefer a tree that summarises the central tree-like trend of evolution. In this paper, we formally investigate methods for transforming an arbitrary phylogenetic network into a tree (on the same set of leaves) and ask which ones (if any) satisfy a simple consistency condition. This consistency condition states that if we add additional species into a phylogenetic network (without otherwise changing this original network) then transforming this enlarged network into a rooted phylogenetic tree induces the same tree on the original set of species as transforming the original network. We show that the LSA (lowest stable ancestor) tree method satisfies this consistency property, whereas several other commonly used methods (and a new one we introduce) do not. We also briefly consider transformations that convert arbitrary phylogenetic networks to another simpler class, namely normal networks.

q-bio.PE

Neutral phylogenetic models and their role in tree-based biodiversity measures

A wide variety of stochastic models of cladogenesis (based on speciation and extinction) lead to an identical distribution on phylogenetic tree shapes once the edge lengths are ignored. By contrast, the distribution of the tree's edge lengths is generally quite sensitive to the underlying model. In this paper, we review the impact of different model choices on tree shape and edge length distribution, and its impact for studying the properties of phylogenetic diversity (PD) as a measure of biodiversity, and the loss of PD as species become extinct at the present. We also compare PD with a stochastic model of feature diversity, and investigate some mathematical links and inequalities between these two measures plus their predictions concerning the loss of biodiversity under extinction at the present.

q-bio.PE

0-1 laws for pattern occurrences in phylogenetic trees and networks

In a recent paper, the question of determining the fraction of binary trees that contain a fixed pattern known as the snowflake was posed. We show that this fraction goes to 1, providing two very different proofs: a purely combinatorial one that is quantitative and specific to this problem; and a proof using branching process techniques that is less explicit, but also much more general, as it applies to any fixed patterns and can be extended to other trees and networks. In particular, it follows immediately from our second proof that the fraction of $d$-ary trees (resp. level-$k$ networks) that contain a fixed $d$-ary tree (resp. level-$k$ network) tends to $1$ as the number of leaves grows.

q-bio.PE

Phylogenetic network classes through the lens of expanding covers

It was recently shown that a large class of phylogenetic networks, the `labellable' networks, is in bijection with the set of `expanding' covers of finite sets. In this paper, we show how several prominent classes of phylogenetic networks can be characterised purely in terms of properties of their associated covers. These classes include the tree-based, tree-child, orchard, tree-sibling, and normal networks.

q-bio.PE

Interior operators and their relationship to autocatalytic networks

The emergence of an autocatalytic network from an available set of elements is a fundamental step in early evolutionary processes, such as the origin of metabolism. Given a set of elements, the reactions between them (chemical or otherwise), and certain elements catalysing certain reactions, a Reflexively Autocatalytic F-generated (RAF) set is a subset $R'$ of reactions that is self-generating from a given food set, and with each reaction in $R'$ being catalysed from within $R'$. RAF theory has been applied to various phenomena in theoretical biology, and a key feature of the approach is that it is possible to efficiently identify and classify RAFs within large systems. This is possible because RAFs can be described as the (nonempty) subsets of the reactions that are the fixed points of an (efficiently computable) interior map that operates on subsets of reactions. Although the main generic results concerning RAFs can be derived using just this property, we show that for systems with at least 12 reactions there are generic results concerning RAFs that cannot be proven using the interior operator property alone.

q-bio.MN

Counting and optimising maximum phylogenetic diversity sets

In conservation biology, phylogenetic diversity (PD) provides a way to quantify the impact of the current rapid extinction of species on the evolutionary `Tree of Life'. This approach recognises that extinction not only removes species but also the branches of the tree on which unique features shared by the extinct species arose. In this paper, we investigate three questions that are relevant to PD. The first asks how many sets of species of given size $k$ preserve the maximum possible amount of PD in a given tree. The number of such maximum PD sets can be very large, even for moderate-sized phylogenies. We provide a combinatorial characterisation of maximum PD sets, focusing on the setting where the branch lengths are ultrametric (e.g. proportional to time). This leads to a polynomial-time algorithm for calculating the number of maximum PD sets of size $k$ by applying a generating function; we also investigate the types of tree shapes that harbour the most (or fewest) maximum PD sets of size $k$. Our second question concerns optimising a linear function on the species (regarded as leaves of the phylogenetic tree) across all the maximum PD sets of a given size. Using the characterisation result from the first question, we show how this optimisation problem can be solved in polynomial time, even though the number of maximum PD sets can grow exponentially. Our third question considers a dual problem: If $k$ species were to become extinct, then what is the largest possible {\em loss} of PD in the resulting tree? For this question, we describe a polynomial-time solution based on dynamical programming.

q-bio.PE

Combinatorics of polymer models of early metabolism

Polymer models are a widely used tool to study the prebiotic formation of metabolism at the origins of life. Counts of the number of reactions in these models are often crucial in probabilistic arguments concerning the emergence of autocatalytic networks. In the first part of this paper, we provide the first exact description of the number of reactions under widely applied model assumptions. Conclusions from earlier studies rely on either approximations or asymptotic counting, and we show that the exact counts lead to similar, though not always identical, asymptotic results. In the second part of the paper, we investigate a novel model assumption whereby polymers are invariant under spatial rotation. We outline the biochemical relevance of this condition and again give exact enumerative and asymptotic formulae for the number of reactions.

q-bio.MN

Defining phylogenetic networks using ancestral profiles

Rooted phylogenetic networks provide a more complete representation of the ancestral relationship between species than phylogenetic trees when reticulate evolutionary processes are at play. One way to reconstruct a phylogenetic network is to consider its `ancestral profile' (the number of paths from each ancestral vertex to each leaf). In general, this information does not uniquely determine the underlying phylogenetic network. A recent paper considered a new class of phylogenetic networks called `orchard networks' where this uniqueness was claimed to hold. Here we show that an additional restriction on the network, that of being `stack-free', is required in order for the original uniqueness claim to hold. On the other hand, if the additional stack-free restriction is lifted, we establish an alternative result; namely, there is uniqueness within the class of orchard networks up to the resolution of vertices of high in-degree.

math.CO

Modelling aspects of consciousness: a topological perspective

Attention Schema Theory (AST) is a recent proposal to provide a scientific explanation for the basis of subjective awareness. In AST, the brain constructs a representation of attention taking place in its own (and others') mind (`the attention schema'). Moreover, this representation is incomplete for efficiency reasons. This inherent incompleteness of the attention schema results in the inability of humans to understand how their own subjective awareness arises (related to the so-called `hard problem' of consciousness). Given this theory, the present paper asks whether a mind (either human or machine-based) that incorporates attention, and that contains a representation of its own attention, can ever have a complete representation. Using a simple yet general model and a mathematical argument based on classical topology, we show that a complete representation of attention is not possible, since it cannot faithfully represent streams of attention. In this way, the study supports one of the core aspects of AST, that the brain's representation of its own attention is necessarily incomplete.

q-bio.NC