SearcharxivSearch

arXiv subjects

Michael M. Desai

Publications and source records attributed to Michael M. Desai.

17 recordsLinked to original sources

Inferring genotype-phenotype maps using attention models

Predicting phenotype from genotype is a central challenge in genetics. Traditional approaches in quantitative genetics typically analyze this problem using methods based on linear regression. These methods generally assume that the genetic architecture of complex traits can be parameterized in terms of an additive model, where the effects of loci are independent, plus (in some cases) pairwise epistatic interactions between loci. However, these models struggle to analyze more complex patterns of epistasis or subtle gene-environment interactions. Recent advances in machine learning, particularly attention-based models, offer a promising alternative. Initially developed for natural language processing, attention-based models excel at capturing context-dependent interactions and have shown exceptional performance in predicting protein structure and function. Here, we apply attention-based models to quantitative genetics. We analyze the performance of this attention-based approach in predicting phenotype from genotype using simulated data across a range of models with increasing epistatic complexity, and using experimental data from a recent quantitative trait locus mapping study in budding yeast. We find that our model demonstrates superior out-of-sample predictions in epistatic regimes compared to standard methods. We also explore a more general multi-environment attention-based model to jointly analyze genotype-phenotype maps across multiple environments and show that such architectures can be used for "transfer learning" - predicting phenotypes in novel environments with limited training data.

q-bio.GN

Social network structure and the spread of complex contagions from a population genetics perspective

Ideas, behaviors, and opinions spread through social networks. If the probability of spreading to a new individual is a non-linear function of the fraction of the individuals' affected neighbors, such a spreading process becomes a "complex contagion". This non-linearity does not typically appear with physically spreading infections, but instead can emerge when the concept that is spreading is subject to game theoretical considerations (e.g. for choices of strategy or behavior) or psychological effects such as social reinforcement and other forms of peer influence (e.g. for ideas, preferences, or opinions). Here we study how the stochastic dynamics of such complex contagions are affected by the underlying network structure. Motivated by simulations of complex epidemics on real social networks, we present a general framework for analyzing the statistics of contagions with arbitrary non-linear adoption probabilities based on the mathematical tools of population genetics. Our framework provides a unified approach that illustrates intuitively several key properties of complex contagions: stronger community structure and network sparsity can significantly enhance the spread, while broad degree distributions dampen the effect of selection. Finally, we show that some structural features can exhibit critical values that demarcate regimes where global epidemics become possible for networks of arbitrary size. Our results draw parallels between the competition of genes in a population and memes in a world of minds and ideas. Our tools provide insight into the spread of information, behaviors, and ideas via social influence, and highlight the role of macroscopic network structure in determining their fate.

physics.soc-ph

The competition between simple and complex evolutionary trajectories in asexual populations

On rugged fitness landscapes where sign epistasis is common, adaptation can often involve either individually beneficial "uphill" mutations or more complex mutational trajectories involving fitness valleys or plateaus. The dynamics of the evolutionary process determine the probability that evolution will take any specific path among a variety of competing possible trajectories. Understanding this evolutionary choice is essential if we are to understand the outcomes and predictability of adaptation on rugged landscapes. We present a simple model to analyze the probability that evolution will eschew immediately uphill paths in favor of crossing fitness valleys or plateaus that lead to higher fitness but less accessible genotypes. We calculate how this probability depends on the population size, mutation rates, and relevant selection pressures, and compare our analytical results to Wright-Fisher simulations. We find that the probability of valley crossing depends nonmonotonically on population size: intermediate size populations are most likely to follow a "greedy" strategy of acquiring immediately beneficial mutations even if they lead to evolutionary dead ends, while larger and smaller populations are more likely to cross fitness valleys to reach distant advantageous genotypes. We explicitly identify the boundaries between these different regimes in terms of the relevant evolutionary parameters. Above a certain threshold population size, we show that the degree of evolutionary "foresight" depends only on a single simple combination of the relevant parameters.

q-bio.PE

The impact of macroscopic epistasis on long-term evolutionary dynamics

Genetic interactions can strongly influence the fitness effects of individual mutations, yet the impact of these epistatic interactions on evolutionary dynamics remains poorly understood. Here we investigate the evolutionary role of epistasis over 50,000 generations in a well-studied laboratory evolution experiment in E. coli. The extensive duration of this experiment provides a unique window into the effects of epistasis during long-term adaptation to a constant environment. Guided by analytical results in the weak-mutation limit, we develop a computational framework to assess the compatibility of a given epistatic model with the observed patterns of fitness gain and mutation accumulation through time. We find that a decelerating fitness trajectory alone provides little power to distinguish between competing models, including those that lack any direct epistatic interactions between mutations. However, when combined with the mutation trajectory, these observables place strong constraints on the set of possible models of epistasis, ruling out many existing explanations of the data. Instead, we find that the data are consistent with "two-epoch" model of adaptation, in which an initial burst of diminishing returns epistasis is followed by a steady accumulation of mutations under a constant distribution of fitness effects. Our results highlight the need for additional DNA sequencing of these populations, as well as for more sophisticated models of epistasis that are compatible with all of the experimental data.

q-bio.PE

Pleiotropic consequences of adaptation across gradations of environmental stress in budding yeast

Adaptation to one environment often results in fitness gains and losses in other conditions. To characterize how these consequences of adaptation depend on the physical similarity between environments, we evolved 180 populations of Saccharomyces cerevisiae at different degrees of stress induced by either salt, temperature, pH, or glucose depletion. We measure how the fitness of clones adapted to each environment depends on the intensity of the corresponding type of stress. We find that clones evolved in a given type and intensity of stress tend to gain fitness in other similar intensities of that stress, and lose progressively more fitness in more physically dissimilar environments. These fitness trade-offs are asymmetric: adaptation to permissive conditions incurred a smaller trade-off in stressful conditions than vice versa. We also find that fitnesses of clones are highly correlated across similar intensities of stress, but these correlations decay towards zero in more dissimilar environments. To interpret these results, we introduce the concept of a joint distribution of fitness effects of new mutations in multiple environments (the JDFE), which describes the probability that a given mutation has particular fitness effects in some set of conditions. We find that our observations are consistent with JDFEs that are highly correlated between physically similar environments, and that become less correlated and asymmetric as the environments become more dissimilar. The JDFE provides a framework for quantifying evolutionary similarity between conditions, and forms a useful basis for theoretical work aimed at predicting the outcomes of evolution in fluctuating environments.

q-bio.PE

The Fates of Mutant Lineages and the Distribution of Fitness Effects of Beneficial Mutations in Laboratory Budding Yeast Populations

The outcomes of evolution are determined by which mutations occur and fix. In rapidly adapting microbial populations, this process is particularly hard to predict because lineages with different beneficial mutations often spread simultaneously and interfere with one another's fixation. Hence to predict the fate of any individual variant, we must know the rate at which new mutations create competing lineages of higher fitness. Here, we directly measured the effect of this interference on the fates of specific adaptive variants in laboratory Saccharomyces cerevisiae populations and used these measurements to infer the distribution of fitness effects of new beneficial mutations. To do so, we seeded marked lineages with different fitness advantages into replicate populations and tracked their subsequent frequencies for hundreds of generations. Our results illustrate the transition between strongly advantageous lineages which decisively sweep to fixation and more moderately advantageous lineages that are often outcompeted by new mutations arising during the course of the experiment. We developed an approximate likelihood framework to compare our data to simulations and found that the effects of these competing beneficial mutations were best approximated by an exponential distribution, rather than one with a single effect size. We then used this inferred distribution of fitness effects to predict the rate of adaptation in a set of independent control populations. Finally, we discuss how our experimental design can serve as a screen for rare, large-effect beneficial mutations.

q-bio.PE

Spatial population expansion promotes the evolution of cooperation in an experimental Prisoner's Dilemma

Cooperation is ubiquitous in nature, but explaining its existence remains a central interdisciplinary challenge. Cooperation is most difficult to explain in the Prisoner's Dilemma game, where cooperators always lose in direct competition with defectors despite increasing mean fitness. Here we demonstrate how spatial population expansion, a widespread natural phenomenon, promotes the evolution of cooperation. We engineer an experimental Prisoner's Dilemma game in the budding yeast Saccharomyces cerevisiae to show that, despite losing to defectors in nonexpanding conditions, cooperators increase in frequency in spatially expanding populations. Fluorescently labeled colonies show genetic demixing of cooperators and defectors, followed by increase in cooperator frequency as cooperator sectors overtake neighboring defector sectors. Together with lattice-based spatial simulations, our results suggest that spatial population expansion drives the evolution of cooperation by (1) increasing positive genetic assortment at population frontiers and (2) selecting for phenotypes maximizing local deme productivity. Spatial expansion thus creates a selective force whereby cooperator-enriched demes overtake neighboring defector-enriched demes in a "survival of the fastest". We conclude that colony growth alone can promote cooperation and prevent defection in microbes. Our results extend to other species with spatially restricted dispersal undergoing range expansion, including pathogens, invasive species, and humans.

q-bio.PE

Interference limits resolution of selection pressures from linked neutral diversity

Pervasive natural selection can strongly influence observed patterns of genetic variation, but these effects remain poorly understood when multiple selected variants segregate in nearby regions of the genome. Classical population genetics fails to account for interference between linked mutations, which grows increasingly severe as the density of selected polymorphisms increases. Here, we describe a simple limit that emerges when interference is common, in which the fitness effects of individual mutations play a relatively minor role. Instead, molecular evolution is determined by the variance in fitness within the population, defined over an effectively asexual segment of the genome (a ``linkage block''). We exploit this insensitivity in a new ``coarse-grained'' coalescent framework, which approximates the effects of many weakly selected mutations with a smaller number of strongly selected mutations with the same variance in fitness. This approximation generates accurate and efficient predictions for the genetic diversity that cannot be summarized by a simple reduction in effective population size. However, these results suggest a fundamental limit on our ability to resolve individual selection pressures from contemporary sequence data alone, since a wide range of parameters yield nearly identical patterns of sequence variability.

q-bio.PE

Fluctuations in fitness distributions and the effects of weak linked selection on sequence evolution

Evolutionary dynamics and patterns of molecular evolution are strongly influenced by selection on linked regions of the genome, but our quantitative understanding of these effects remains incomplete. Recent work has focused on predicting the distribution of fitness within an evolving population, and this forms the basis for several methods that leverage the fitness distribution to predict the patterns of genetic diversity when selection is strong. However, in weakly selected populations random fluctuations due to genetic drift are more severe, and neither the distribution of fitness nor the sequence diversity within the population are well understood. Here, we briefly review the motivations behind the fitness-distribution picture, and summarize the general approaches that have been used to analyze this distribution in the strong-selection regime. We then extend these approaches to the case of weak selection, by outlining a perturbative treatment of selection at a large number of linked sites. This allows us to quantify the stochastic behavior of the fitness distribution and yields exact analytical predictions for the sequence diversity and substitution rate in the limit that selection is weak.

q-bio.PE

Genetic Diversity and the Structure of Genealogies in Rapidly Adapting Populations

Positive selection distorts the structure of genealogies and hence alters patterns of genetic variation within a population. Most analyses of these distortions focus on the signatures of hitchhiking due to hard or soft selective sweeps at a single genetic locus. However, in linked regions of rapidly adapting genomes, multiple beneficial mutations at different loci can segregate simultaneously within the population, an effect known as clonal interference. This leads to a subtle interplay between hitchhiking and interference effects, which leads to a unique signature of rapid adaptation on genetic variation both at the selected sites and at linked neutral loci. Here, we introduce an effective coalescent theory (a "fitness-class coalescent") that describes how positive selection at many perfectly linked sites alters the structure of genealogies. We use this theory to calculate several simple statistics describing genetic variation within a rapidly adapting population, and to implement efficient backwards-time coalescent simulations which can be used to predict how clonal interference alters the expected patterns of molecular evolution.

q-bio.PE

Rare beneficial mutations can halt Muller's ratchet

The vast majority of mutations are deleterious, and are eliminated by purifying selection. Yet in finite asexual populations, purifying selection cannot completely prevent the accumulation of deleterious mutations due to Muller's ratchet: once lost by stochastic drift, the most-fit class of genotypes is lost forever. If deleterious mutations are weakly selected, Muller's ratchet turns into a mutational "meltdown" leading to a rapid degradation of population fitness. Evidently, the long term stability of an asexual population requires an influx of beneficial mutations that continuously compensate for the accumulation of the weakly deleterious ones. Here we propose that the stable evolutionary state of a population in a static environment is a dynamic mutation-selection balance, where accumulation of deleterious mutations is on average offset by the influx of beneficial mutations. We argue that this state exists for any population size N and mutation rate $U$. Assuming that beneficial and deleterious mutations have the same fitness effect s, we calculate the fraction of beneficial mutations, ε, that maintains the balanced state. We find that a surprisingly low εsuffices to maintain stability, even in small populations in the face of high mutation rates and weak selection. This may explain the maintenance of mitochondria and other asexual genomes, and has implications for the expected statistics of genetic diversity in these populations.

q-bio.PE

The structure of allelic diversity in the presence of purifying selection

In the absence of selection, the structure of allelic diversity is described by the elegant sampling formula of Ewens. This formula has helped shape our expectations of empirical patterns of molecular variation. Along with coalescent theory, it provides statistical techniques for rejecting the null model of neutrality. However, we still do not fully understand the statistics of the allelic diversity we expect to see in the presence of natural selection. Earlier work has described the effects of strongly deleterious mutations linked to many neutral sites, and allelic variation in models where offspring fitness is unrelated to parental fitness, but it has proven difficult to understand allelic diversity in the presence of purifying selection at many linked sites. Here, we study the population genetics of infinitely many perfectly linked sites, some neutral and some deleterious. Our approach is based on studying the lineage structure within each class of individuals of similar fitness in the deleterious mutation-selection balance. Analogous to the Ewens sampling formula, we derive expressions for the likelihoods of any configuration of allelic types in a sample. We find that for moderate and weak selection pressures the patterns of allelic diversity cannot be described by a neutral model for any choice of the effective population size, indicating that there is power to detect selection from patterns of sampled allelic diversity.

q-bio.PE

The Structure of Genealogies in the Presence of Purifying Selection: A "Fitness-Class Coalescent"

Compared to a neutral model, purifying selection distorts the structure of genealogies and hence alters the patterns of sampled genetic variation. Although these distortions may be common in nature, our understanding of how we expect purifying selection to affect patterns of molecular variation remains incomplete. Genealogical approaches such as coalescent theory have proven difficult to generalize to situations involving selection at many linked sites, unless selection pressures are extremely strong. Here, we introduce an effective coalescent theory (a "fitness-class coalescent") to describe the structure of genealogies in the presence of purifying selection at many linked sites. We use this effective theory to calculate several simple statistics describing the expected patterns of variation in sequence data, both at the sites under selection and at linked neutral sites. Our analysis combines our earlier description of the allele frequency spectrum in the presence of purifying selection (Desai et al. 2010) with the structured coalescent approach of Nordborg (1997), to trace the ancestry of individuals through the distribution of fitnesses within the population. Alternatively, we can derive our results using an extension of the coalescent approach of Hudson and Kaplan (1994). We find that purifying selection leads to patterns of genetic variation that are related but not identical to a neutrally evolving population in which population size has varied in a specific way in the past.

q-bio.PE

Clonal Interference, Multiple Mutations, and Adaptation in Large Asexual Populations

Two important problems affect the ability of asexual populations to accumulate beneficial mutations, and hence to adapt. First, clonal interference causes some beneficial mutations to be outcompeted by more-fit mutations which occur in the same genetic background. Second, multiple mutations occur in some individuals, so even mutations of large effect can be outcompeted unless they occur in a good genetic background which contains other beneficial mutations. In this paper, we use a Monte Carlo simulation to study how these two factors influence the adaptation of asexual populations. We find that the results depend qualitatively on the shape of the distribution of the effects of possible beneficial mutations. When this distribution falls off slower than exponentially, clonal interference alone reasonably describes which mutations dominate the adaptation, although it gives a misleading picture of the evolutionary dynamics. When the distribution falls off faster than exponentially, an analysis based on multiple mutations is more appropriate. Using our simulations, we are able to explore the limits of validity of both of these approaches, and we explore the complex dynamics in the regimes where neither are fully applicable.

q-bio.PE

Beneficial mutation-selection balance and the effect of linkage on positive selection

When beneficial mutations are rare, they accumulate by a series of selective sweeps. But when they are common, many beneficial mutations will occur before any can fix, so there will be many different mutant lineages in the population concurrently. In an asexual population, these different mutant lineages interfere and not all can fix simultaneously. In addition, further beneficial mutations can accumulate in mutant lineages while these are still a minority of the population. In this paper, we analyze the dynamics of such multiple mutations and the interplay between multiple mutations and interference between clones. These result in substantial variation in fitness accumulating within a single asexual population. The amount of variation is determined by a balance between selection, which destroys variation, and beneficial mutations, which create more. The behavior depends in a subtle way on the population parameters: the population size, the beneficial mutation rate, and the distribution of the fitness increments of the potential beneficial mutations. The mutation-selection balance leads to a continually evolving population with a steady-state fitness variation. This variation increases logarithmically with both population size and mutation rate and sets the rate at which the population accumulates beneficial mutations, which thus also grows only logarithmically with population size and mutation rate. These results imply that mutator phenotypes are less effective in larger asexual populations. They also have consequences for the advantages (or disadvantages) of sex via the Fisher-Muller effect; these are discussed briefly.

q-bio.PE

Synonymous codon usage and selection on proteins

Selection pressures on proteins are usually measured by comparing homologous nucleotide sequences (Zuckerkandl and Pauling 1965). Recently we introduced a novel method, termed `volatility', to estimate selection pressures on protein sequences from their synonymous codon usage (Plotkin and Dushoff 2003, Plotkin et al 2004a). Here we provide a theoretical foundation for this approach. We derive the expected frequencies of synonymous codons as a function of the strength of selection, the mutation rate, and the effective population size. We analyze the conditions under which we can expect to draw inferences from biased codon usage, and we estimate the time scales required to establish and maintain such a signal. Our results indicate that, over a broad range of parameters, synonymous codon usage can reliably distinguish between negative selection, positive selection, and neutrality. While the power of volatility to detect negative selection depends on the population size, there is no such dependence for the detection of positive selection. Furthermore, we show that phenomena such as transient hyper-mutators in microbes can improve the power of volatility to detect negative selection, even when the typical observed neutral site heterozygosity is low.

q-bio.PE

A Quasispecies on a Moving Oasis

A population evolving in an inhomogeneous environment will adapt differently to different regions. We study the conditions under which such a population can maintain adaptations to a particular region when that region is not stationary, but can move. In particular, we study a quasispecies living near a favorable patch ("oasis") in the middle of a large "desert." The population has two genetic states, one of which which conveys a relative advantage while in the oasis at the cost of a disadvantage in the desert. We consider the population dynamics when the oasis is moving, or equivalently some form of "wind" is blowing the population away from the oasis. We find that the ratio of the two types of individuals exhibits sharp transitions at particular oasis velocities. We calculate an extinction velocity, and a switching velocity above which the dominance switches from the oasis-adapted genotype to the desert-adapted one. This switching velocity is analagous to the quasispecies mutational error threshold. Above this velocity, the population cannot maintain adaptations to the properties of the oasis.

q-bio.PE