SearcharxivSearch

arXiv subjects

Amaury Lambert

Publications and source records attributed to Amaury Lambert.

At least 19 recordsLinked to original sources

Geometry and stability of species complexes: larger species speciate less often

Species complexes are groups of closely related populations exchanging genes through dispersal. We study the dynamics of the structure of species complexes in a class of metapopulation models where demes can exchange genetic material through migration and diverge through the accumulation of new mutations. Importantly, we model the ecological feedback of differentiation on gene flow by assuming that the success of migrations decreases with genetic distance, through a specific function $h$. We investigate the effects of metapopulation size on the coherence of species structures, depending on some mathematical characteristics of the feedback function $h$. Our results suggest that with larger metapopulation sizes, species form increasingly coherent, transitive, and uniform entities. We conclude that the initiation of speciation events in large species requires the existence of idiosyncratic geographic or selective restrictions on gene flow.

q-bio.PE

Evolution of a trait distributed over a large fragmented population: Propagation of chaos meets adaptive dynamics

We consider a metapopulation made up of $K$ demes, each containing $N$ individuals bearing a heritable quantitative trait. Demes are connected by migration and undergo independent Moran processes with mutation and selection based on trait values. Mutation and migration rates are tuned so that each deme receives a migrant or a mutant in the same slow timescale and is thus essentially monomorphic at all times for the trait (adaptive dynamics). In the timescale of mutation/migration, the metapopulation can then be seen as a giant spatial Moran model with size $K$ that we characterize. As $K\to \infty$ and physical space becomes continuous, the empirical distribution of the trait (over the physical and trait spaces) evolves deterministically according to an integro-differential evolution equation. In this limit, the trait of every migrant is drawn from this global distribution, so that conditional on its initial state, traits from finitely many demes evolve independently (propagation of chaos). Under mean-field dispersal, the value $X_t$ of the trait at time $t$ and at any given location has a law denoted $\mu_t$ and a jump kernel with two terms: a mutation-fixation term and a migration-fixation term involving $\mu_{t-}$ (McKean-Vlasov equation). In the limit where mutations have small effects and migration is further slowed down accordingly, we obtain the convergence of $X$, in the new migration timescale, to the solution of a stochastic differential equation which can be referred to as a new canonical equation of adaptive dynamics. This equation includes an advection term representing selection, a diffusive term due to genetic drift, and a jump term, representing the effect of migration, to a state distributed according to its own law.

math.PR

The gene's-eye view of quantitative genetics

Modelling the evolution of a continuous trait in a biological population is one of the oldest problems in evolutionary biology, which led to the birth of quantitative genetics. With the recent development of GWAS methods, it has become essential to link the evolution of the trait distribution to the underlying evolution of allelic frequencies at many loci, co-contributing to the trait value. The way most articles go about this is to make assumptions on the trait distribution, and use Wright's formula to model how the evolution of the trait translates on each individual locus. Here, we take a gene's eye-view of the system, starting from an explicit finite-loci model with selection, drift, recombination and mutation, in which the trait value is a direct product of the genome. We let the number of loci go to infinity under the assumption of strong recombination, and characterize the limit behavior of a given locus with a McKean-Vlasov SDE and the corresponding Fokker-Planck IPDE. In words, the selection on a typical locus depends on the mean behaviour of the other loci which can be approximated with the law of the focal locus. Results include the independence of two loci and explicit stationary distribution for allelic frequencies at a given locus (under some assumptions on the fitness function).

math.PR

Unsupervised detection and fitness estimation of emerging SARS-CoV-2 variants. Application to wastewater samples (ANRS0160)

Repeated waves of emerging variants during the SARS-CoV-2 pandemics have highlighted the urge of collecting longitudinal genomic data and developing statistical methods based on time series analyses for detecting new threatening lineages and estimating their fitness early in time. Most models study the evolution of the prevalence of particular lineages over time and require a prior classification of sequences into lineages. Such process is prone to induce delays and bias. More recently, few authors studied the evolution of the prevalence of mutations over time with alternative clustering approaches, avoiding specific lineage classification. Most of the aforementioned methods are however either non parametric or unsuited to pooled data characterizing, for instance, wastewater samples. In this context, we propose an alternative unsupervised method for clustering mutations according to their frequency trajectory over time and estimating group fitness from time series of pooled mutation prevalence data. Our model is a mixture of observed count data and latent group assignment and we use the expectation-maximization algorithm for model selection and parameter estimation. The application of our method to time series of SARS-CoV-2 sequencing data collected from wastewater treatment plants in France from October 2020 to April 2021 shows its ability to agnostically group mutations according to their probability of belonging to B.1.160, Alpha, Beta, B.1.177 variants with selection coefficient estimates per group in coherence with the viral dynamics in France reported by Nextstrain. Moreover, our method detected the Alpha variant as threatening as early as supervised methods (which track specific mutations over time) with the noticeable difference that, since unsupervised, it does not require any prior information on the set of mutations.

q-bio.QM

Ages, sizes and (trees within) trees of taxa and of urns, from Yule to today

The paper written in 1925 by G. Udny Yule that we celebrate in this special issue introduces several novelties and results that we recall in detail. First, we discuss Yule (1925)'s main legacies over the past century, focusing on empirical frequency distributions with heavy tails and random tree models for phylogenies. We estimate the year when Yule's work was re-discovered by scientists interested in stochastic processes of population growth (1948) and the year from which it began to be cited (1951, Yule's death). We highlight overlooked aspects of Yule's work (e.g., the Yule process of Yule processes) and correct some common misattributions (e.g., the Yule tree). Second, we generalize Yule's results on the average frequency of genera of a given age and size (number of species). We show that his formula also applies to the age $A$ and size $S$ of any randomly chosen genus and that the pairs $(A_i, S_i)$ are equally distributed and independent across genera. This property extends to triples $(H_i, A_i, S_i)$, where $(H_i)$ are the coalescence times of the genus phylogeny, even when species diversification within genera follows any integer-valued process, including species extinctions. Studying $(A, S)$ in this broader context allows us to identify cases where $S$ has a power-law tail distribution, with new applications to urn schemes.

q-bio.PE

Neutral Diversity in Experimental Metapopulations

New automated and high-throughput methods allow the manipulation and selection of numerous bacterial populations. In this manuscript we are interested in the neutral diversity patterns that emerge from such a setup in which many bacterial populations are grown in parallel serial transfers, in some cases with population-wide extinction and splitting events. We model bacterial growth by a birth-death process and use the theory of coalescent point processes. We show that there is a dilution factor that optimises the expected amount of neutral diversity for a given amount of cycles, and study the power law behaviour of the mutation frequency spectrum for different experimental regimes. We also explore how neutral variation diverges between two recently split populations by establishing a new formula for the expected number of shared and private mutations. Finally, we show the interest of such a setup to select a phenotype of interest that requires multiple mutations.

q-bio.PE

From individual-based epidemic models to McKendrick-von Foerster PDEs: A guide to modeling and inferring COVID-19 dynamics

We present a unifying, tractable approach for studying the spread of viruses causing complex diseases requiring to be modeled using a large number of types (e.g., infective stage, clinical state, risk factor class). We show that recording each infected individual's infection age, i.e., the time elapsed since infection, has three benefits. First, regardless of the number of types, the age distribution of the population can be described by means of a first-order, one-dimensional partial differential equation (PDE) known as the McKendrick-von Foerster equation. The frequency of type $i$ is simply obtained by integrating the probability of being in state $i$ at a given age against the age distribution. This representation induces a simple methodology based on the additional assumption of Poisson sampling to infer and forecast the epidemic. We illustrate this technique using French data from the COVID-19 epidemic. Second, our approach generalizes and simplifies standard compartmental models using high-dimensional systems of ordinary differential equations (ODEs) to account for disease complexity. We show that such models can always be rewritten in our framework, thus, providing a low-dimensional yet equivalent representation of these complex models. Third, beyond the simplicity of the approach, we show that our population model naturally appears as a universal scaling limit of a large class of fully stochastic individual-based epidemic models, where the initial condition of the PDE emerges as the limiting age structure of an exponentially growing population starting from a single individual.

q-bio.PE

The coalescent structure of uniform and Poisson samples from multitype branching processes

We introduce a Poissonization method to study the coalescent structure of uniform samples from branching processes. This method relies on the simple observation that a uniform sample of size $k$ taken from a random set with positive Lebesgue measure may be represented as a mixture of Poisson samples with rate $λ$ and mixing measure $k \mathrm{d} λ/ λ$. We develop a multitype analogue of this mixture representation, and use it to characterise the coalescent structure of multitype continuous-state branching processes in terms of random multitype forests. Thereafter we study the small time asymptotics of these random forests, establishing a correspondence between multitype continuous-state branching proesses and multitype $Λ$-coalescents.

math.PR

Combinatorial and stochastic properties of ranked tree-child networks

Tree-child networks are a recently-described class of directed acyclic graphs that have risen to prominence in phylogenetics (the study of evolutionary trees and networks). Although these networks have a number of attractive mathematical properties, many combinatorial questions concerning them remain intractable. In this paper, we show that endowing these networks with a biologically relevant ranking structure yields mathematically tractable objects, which we term ranked tree-child networks (RTCNs). We explain how to derive exact and explicit combinatorial results concerning the enumeration and generation of these networks. We also explore probabilistic questions concerning the properties of RTCNs when they are sampled uniformly at random. These questions include the lengths of random walks between the root and leaves (both from the root to the leaves and from a leaf to the root); the distribution of the number of cherries in the network; and sampling RTCNs conditional on displaying a given tree. We also formulate a conjecture regarding the scaling limit of the process that counts the number of lineages in the ancestry of a leaf. The main idea in this paper, namely using ranking as a way to achieve combinatorial tractability, may also extend to other classes of networks.

math.PR

Chromosome Painting: how recombination mixes ancestral colors

We consider a Moran model with recombination in a haploid population of size $N$. At each birth event, with probability $1-ρ_N R$ the offspring copies one parent's chromosome, and with probability $ρ_N R$ she inherits a chromosome that is a mosaic of both parental chromosomes. We assume that at time $0$ each individual has her chromosome painted in a different color and we study the color partition of the chromosome that is asymptotically fixed in a large population, when we look at a portion of the chromosome such that $ρ:= \lim_{N\to \infty} \frac{ρ_N N}{2}\to \infty$. To do so, we follow backwards in time the ancestry of the chromosome of a randomly sampled individual. This yields a Markov process valued in the color partitions of the half-line, that was introduced by Esser et al. (2016), in which blocks can merge and split, called the partitioning process. Its stationary distribution is closely related to the fixed chromosome in our Moran model with recombination. We are able to provide an approximation of this stationary distribution when $ρ\gg 1$ and an error bound. This allows us to show that the distribution of the (renormalised) length of the leftmost block of the partition (i.e. the region of the chromosome that carries the same color as 0) converges to an exponential distribution. In addition, the geometry of this block can be described in terms of a Poisson point process with an explicit intensity measure.

math.PR

Exchangeable coalescents, ultrametric spaces, nested interval-partitions: A unifying approach

Kingman (1978)'s representation theorem states that any exchangeable partition of $\mathbb{N}$ can be represented as a paintbox based on a random mass-partition. Similarly, any exchangeable composition (i.e. ordered partition of $\mathbb{N}$) can be represented as a paintbox based on an interval-partition (Gnedin 1997). Our first main result is that any exchangeable coalescent process (not necessarily Markovian) can be represented as a paintbox based on a random non-decreasing process valued in interval-partitions, called nested interval-partition, generalizing the notion of comb metric space introduced by Lambert & Uribe Bravo (2017) to represent compact ultrametric spaces. As a special case, we show that any $Λ$-coalescent can be obtained from a paintbox based on a unique random nested interval partition called $Λ$-comb, which is Markovian with explicit transitions. This nested interval-partition directly relates to the flow of bridges of Bertoin & Le Gall (2003). We also display a particularly simple description of the so-called evolving coalescent (Pfaffelhuber & Wakolbinger 2006) by a comb-valued Markov process. Next, we prove that any measured ultrametric space $U$, under mild measure-theoretic assumptions on $U$, is the leaf set of a tree composed of a separable subtree called the backbone, on which are grafted additional subtrees, which act as star-trees from the standpoint of sampling. Displaying this so-called weak isometry requires us to extend the Gromov-weak topology of Greven et al (2006), that was initially designed for separable metric spaces, to non-separable ultrametric spaces. It allows us to show that for any such ultrametric space $U$, there is a nested interval-partition which is 1) indistinguishable from $U$ in the Gromov-weak topology; 2) weakly isometric to $U$ if $U$ has complete backbone; 3) isometric to $U$ if $U$ is complete and separable.

math.PR

Kingman's coalescent with erosion

Consider the Markov process taking values in the partitions of N such that each pair of blocks merges at rate one, and each integer is eroded, i.e., becomes a singleton block, at rate d. This is a special case of exchangeable fragmentation-coalescence process called Kingman's coalescent with erosion. We provide a new construction of the stationary distribution of this process as a sample from a standard flow of bridges. This allows us to give a representation of the asymptotic frequencies of this stationary distribution in terms of a sequence of hierarchically independent diffusions. Moreover, we introduce a new process called Kingman's coalescent with immigration, where pairs of blocks coalesce at rate one, and new blocks of size one immigrate at rate d. By coupling Kingman's coalescents with erosion and with immigration, we are able to show that the size of a block chosen uniformly at random from the stationary distribution of the restriction of Kingman's coalescent with erosion to {1,...,n} converges to the total progeny of a critical binary branching process.

math.PR

Trees within trees: Simple nested coalescents

We consider the compact space of pairs of nested partitions of $\mathbb N$, where by analogy with models used in molecular evolution, we call "gene partition" the finer partition and "species partition" the coarser one. We introduce the class of nondecreasing processes valued in nested partitions, assumed Markovian and with exchangeable semigroup. These processes are said simple when each partition only undergoes one coalescence event at a time (but possibly the same time). Simple nested exchangeable coalescent (SNEC) processes can be seen as the extension of $Λ$-coalescents to nested partitions. We characterize the law of SNEC processes as follows. In the absence of gene coalescences, species blocks undergo $Λ$-coalescent type events and in the absence of species coalescences, gene blocks lying in the same species block undergo i.i.d. $Λ$-coalescents. Simultaneous coalescence of the gene and species partitions are governed by an intensity measure $ν_s$ on $(0,1]\times {\mathcal M}_1 ([0,1])$ providing the frequency of species merging and the law in which are drawn (independently) the frequencies of genes merging in each coalescing species block. As an application, we also study the conditions under which a SNEC process comes down from infinity.

math.PR

Coagulation-transport equations and the nested coalescents

The nested Kingman coalescent describes the dynamics of particles (called genes) contained in larger components (called species), where pairs of species coalesce at constant rate and pairs of genes coalesce at constant rate provided they lie within the same species. We prove that starting from $rn$ species, the empirical distribution of species masses (numbers of genes$/n$) at time $t/n$ converges as $n\to\infty$ to a solution of the deterministic coagulation-transport equation $$ \partial_t d \ = \ \partial_x ( ψd ) \ + \ a(t)\left(d\star d - d \right), $$ where $ψ(x) = cx^2$, $\star$ denotes convolution and $a(t)= 1/(t+δ)$ with $δ=2/r$. The most interesting case when $δ=0$ corresponds to an infinite initial number of species. This equation describes the evolution of the distribution of species of mass $x$, where pairs of species can coalesce and each species' mass evolves like $\dot x = -ψ(x)$. We provide two natural probabilistic solutions of the latter IPDE and address in detail the case when $δ=0$. The first solution is expressed in terms of a branching particle system where particles carry masses behaving as independent continuous-state branching processes. The second one is the law of the solution to the following McKean-Vlasov equation $$ dx_t \ = \ - ψ(x_t) \,dt \ + \ v_t\,ΔJ_t $$ where $J$ is an inhomogeneous Poisson process with rate $1/(t+δ)$ and $(v_t; t\geq0)$ is a sequence of independent rvs such that $ {\mathcal L}(v_t) = {\mathcal L}(x_t)$. We show that there is a unique solution to this equation and we construct this solution with the help of a marked Brownian coalescent point process. When $ψ(x)=x^γ$, we show the existence of a self-similar solution for the PDE which relates when $γ=2$ to the speed of coming down from infinity of the nested Kingman coalescent.

math.PR

The split-and-drift random graph, a null model for speciation

We introduce a new random graph model motivated by biological questions relating to speciation. This random graph is defined as the stationary distribution of a Markov chain on the space of graphs on $\{1, \ldots, n\}$. The dynamics of this Markov chain is governed by two types of events: vertex duplication, where at constant rate a pair of vertices is sampled uniformly and one of these vertices loses its incident edges and is rewired to the other vertex and its neighbors; and edge removal, where each edge disappears at constant rate. Besides the number of vertices $n$, the model has a single parameter $r_n$. Using a coalescent approach, we obtain explicit formulas for the first moments of several graph invariants such as the number of edges or the number of complete subgraphs of order $k$. These are then used to identify five non-trivial regimes depending on the asymptotics of the parameter $r_n$. We derive an explicit expression for the degree distribution, and show that under appropriate rescaling it converges to classical distributions when the number of vertices goes to infinity. Finally, we give asymptotic bounds for the number of connected components, and show that in the sparse regime the number of edges is Poissonian.

math.PR

Totally Ordered Measured Trees and Splitting Trees with Infinite Variation II: Prolific Skeleton Decomposition

The first part of this paper ( arXiv:1607.02114 ) introduced splitting trees, those chronological trees admitting the self-similarity property where individuals give birth, at constant rate, to iid copies of themselves. It also established the intimate relationship between splitting trees and L\'evy processes. The chronological trees involved were formalized as Totally Ordered Measured (TOM) trees. The aim of this paper is to continue this line of research in two directions: we first decompose locally compact TOM trees in terms of their prolific skeleton (consisting of its infinite lines of descent). When applied to splitting trees, this implies the construction of the supercritical ones (which are locally compact) in terms of the subcritical ones (which are compact) grafted onto a Yule tree (which corresponds to the prolific skeleton). As a second (related) direction, we study the genealogical tree associated to our chronological construction. This is done through the technology of the height process introduced by Duquesne and Le Gall. In particular we prove a Ray-Knight type theorem which extends the one for (sub)critical L\'evy trees to the supercritical case.

math.PR

The sequential loss of allelic diversity

This paper gives a new flavor of what Peter Jagers and his co-authors call `the path to extinction'. In a neutral population with constant size $N$, we assume that each individual at time $0$ carries a distinct type, or allele. We consider the joint dynamics of these $N$ alleles, for example the dynamics of their respective frequencies and more plainly the nonincreasing process counting the number of alleles remaining by time $t$. We call this process the extinction process. We show that in the Moran model, the extinction process is distributed as the process counting (in backward time) the number of common ancestors to the whole population, also known as the block counting process of the $N$-Kingman coalescent. Stimulated by this result, we investigate: (1) whether it extends to an identity between the frequencies of blocks in the Kingman coalescent and the frequencies of alleles in the extinction process, both evaluated at jump times; (2) whether it extends to the general case of $\Lambda$-Fleming-Viot processes.

math.PR

Pathogen evolution: slow and steady spreads the best

The theory of life history evolution provides a powerful framework to understand the evolutionary dynamics of pathogens in both epidemic and endemic situations. This framework, however, relies on the assumption that pathogen populations are very large and that one can neglect the effects of demographic stochasticity. Here we expand the theory of life history evolution to account for the effects of finite population size on the evolution of pathogen virulence. We show that demographic stochasticity introduces additional evolutionary forces that can qualitatively affect the dynamics and the evolutionary outcome. We discuss the importance of the shape of pathogen fitness landscape and host heterogeneity on the balance between mutation, selection and genetic drift. In particular, we discuss scenarios where finite population size can dramatically affect classical predictions of deterministic models. This analysis reconciles Adaptive Dynamics with population genetics in finite populations and thus provides a new theoretical toolbox to study life-history evolution in realistic ecological scenarios.

q-bio.PE