SearcharxivSearch

arXiv subjects

Anne-Florence Bitbol

Publications and source records attributed to Anne-Florence Bitbol.

At least 19 recordsLinked to original sources

Out-of-equilibrium selection pressure enhances inference from protein sequence data

Homologous proteins have similar three-dimensional structures and biological functions that shape their sequences. The resulting coevolution-driven correlations underlie methods from Potts models to AlphaFold, which infer protein structure and function from sequences. Using a minimal model, we show that fluctuating selection strength and the onset of new selection pressures improve coevolution-based inference of structural contacts. Our conclusions extend to realistic synthetic data and to the inference of interaction partners. Out-of-equilibrium noise arising from ubiquitous variations in natural selection thus enhances, rather than hinders, the success of inference from protein sequences.

q-bio.BM

Promotion of cooperation in deme-structured populations with growth-merging dynamics

The spatial structure of populations may promote the emergence and maintenance of cooperation. Cooperation in the prisoner's dilemma is favored under specific update rules in evolutionary graph theory models with one individual per node of a graph, but this effect vanishes in models with well-mixed demes connected by migrations under soft selection (each deme contributing independently of its productivity). In contrast, experiments and models involving cycles of growth, merging and dilution have shown that spatial structure can favor cooperation. Here, we reconcile these findings by studying deme-structured populations under growth-merging-dilution dynamics, corresponding to a clique (fully connected graph) under hard selection (each deme contributing in proportion to its productivity). We obtain analytical conditions for the cooperator fraction to increase during deterministic logistic growth, and to increase on average under dilution-growth-merging cycles, in the weak selection regime, where fitness differences between genotypes are small. Furthermore, we analytically express the fixation probability of cooperators under weak selection, yielding a criterion for cooperators to have a higher fixation probability than neutral mutants. Finally, numerical simulations show that stochastic growth further promotes cooperation. Overall, hard selection is essential for cooperation to be promoted in deme-structured populations.

q-bio.PE

Environment heterogeneity creates fast amplifiers of natural selection in graph-structured populations

Complex spatial structure, with partially isolated subpopulations, and environment heterogeneity, such as gradients in nutrients, oxygen, and drugs, both shape the evolution of natural populations. We investigate the impact of environment heterogeneity on mutant fixation in spatially structured populations with demes on the nodes of a graph. When migrations between demes are frequent, we find that environment heterogeneity can amplify natural selection and simultaneously accelerate mutant fixation and extinction, thereby fostering the quick fixation of beneficial mutants. We demonstrate this effect in the star graph, and more strongly in the line graph. We show that amplification requires mutants to have a stronger fitness advantage in demes with stronger migration outflow, and that this condition allows amplification in more general graphs. As a baseline, we consider circulation graphs, where migration inflow and outflow are equal in each deme. In this case, environment heterogeneity has no impact to first order, but increases the fixation probability of beneficial mutants to second order. Finally, when migrations between demes are rare, we show that environment heterogeneity can also foster amplification of selection, by allowing demes with sufficient mutant advantage to become refugia for mutants.

q-bio.PE

Impact of complex spatial population structure on early and long-term adaptation in rugged fitness landscapes

We investigate the exploration of rugged fitness landscapes by spatially structured populations with demes on the nodes of a graph, connected by migrations. In the rare migration regime, we find that finite structures can adapt more efficiently than very large ones, especially in high-dimensional fitness landscapes. Furthermore, we show that, in most landscapes, migration asymmetries associated with some suppression of natural selection allow the population to reach higher fitness peaks first. In this sense, suppression of selection can make early adaptation more efficient. However, the time it takes to reach the first fitness peak is then increased. We also find that suppression of selection tends to enhance finite-size effects. We extend our study to frequent migrations, suggesting that our conclusions hold in this regime. We then investigate the impact of spatial structure with rare migrations on long-term evolution by studying the steady state of the population. For this, we define an effective population size for the steady-state distribution. We find that suppression of selection is associated to reduced steady-state effective population sizes, and reduced average steady-state fitnesses.

q-bio.PE

DiffPaSS -- High-performance differentiable pairing of protein sequences using soft scores

Identifying interacting partners from two sets of protein sequences has important applications in computational biology. Interacting partners share similarities across species due to their common evolutionary history, and feature correlations in amino acid usage due to the need to maintain complementary interaction interfaces. Thus, the problem of finding interacting pairs can be formulated as searching for a pairing of sequences that maximizes a sequence similarity or a coevolution score. Several methods have been developed to address this problem, applying different approximate optimization methods to different scores. We introduce DiffPaSS, a differentiable framework for flexible, fast, and hyperparameter-free optimization for pairing interacting biological sequences, which can be applied to a wide variety of scores. We apply it to a benchmark prokaryotic dataset, using mutual information and neighbor graph alignment scores. DiffPaSS outperforms existing algorithms for optimizing the same scores. We demonstrate the usefulness of our paired alignments for the prediction of protein complex structure. DiffPaSS does not require sequences to be aligned, and we also apply it to non-aligned sequences from T cell receptors.

q-bio.BM

Spatial structure facilitates evolutionary rescue by drug resistance

Bacterial populations often have complex spatial structures, which can impact their evolution. Here, we study how spatial structure affects the evolution of antibiotic resistance in a bacterial population. We consider a minimal model of spatially structured populations where all demes (i.e., subpopulations) are identical and connected to each other by identical migration rates. We show that spatial structure can facilitate the survival of a bacterial population to antibiotic treatment, starting from a sensitive inoculum. Specifically, the bacterial population can be rescued if antibiotic resistant mutants appear and are present when drug is added, and spatial structure can impact the fate of these mutants and the probability that they are present. Indeed, the probability of fixation of neutral or deleterious mutations providing drug resistance is increased in smaller populations. This promotes local fixation of resistant mutants in the structured population, which facilitates evolutionary rescue by drug resistance in the rare mutation regime. Once the population is rescued by resistance, migrations allow resistant mutants to spread in all demes. Our main result that spatial structure facilitates evolutionary rescue by antibiotic resistance extends to more complex spatial structures, and to the case where there are resistant mutants in the inoculum.

q-bio.PE

Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants

AI assistants are being increasingly used by students enrolled in higher education institutions. While these tools provide opportunities for improved teaching and education, they also pose significant challenges for assessment and learning outcomes. We conceptualize these challenges through the lens of vulnerability, the potential for university assessments and learning outcomes to be impacted by student use of generative AI. We investigate the potential scale of this vulnerability by measuring the degree to which AI assistants can complete assessment questions in standard university-level STEM courses. Specifically, we compile a novel dataset of textual assessment questions from 50 courses at EPFL and evaluate whether two AI assistants, GPT-3.5 and GPT-4 can adequately answer these questions. We use eight prompting strategies to produce responses and find that GPT-4 answers an average of 65.8% of questions correctly, and can even produce the correct answer across at least one prompting strategy for 85.1% of questions. When grouping courses in our dataset by degree program, these systems already pass non-project assessments of large numbers of core courses in various degree programs, posing risks to higher education accreditation that will be amplified as these models improve. Our results call for revising program-level assessment design in higher education in light of advances in generative AI.

cs.CY

Bridging Wright-Fisher and Moran models

The Wright-Fisher model and the Moran model are both widely used in population genetics. They describe the time evolution of the frequency of an allele in a well-mixed population with fixed size. We propose a simple and tractable model which bridges the Wright-Fisher and the Moran descriptions. We assume that a fixed fraction of the population is updated at each discrete time step. In this model, we determine the fixation probability of a mutant and its average fixation and extinction times, under the diffusion approximation. We further study the associated coalescent process, which converges to Kingman's coalescent, and we calculate effective population sizes. We generalize our model, first by taking into account fluctuating updated fractions or individual lifetimes, and then by incorporating selection on the lifetime as well as on the reproductive fitness.

q-bio.PE

Impact of phylogeny on the inference of functional sectors from protein sequence data

Statistical analysis of multiple sequence alignments of homologous proteins has revealed groups of coevolving amino acids called sectors. These groups of amino-acid sites feature collective correlations in their amino-acid usage, and they are associated to functional properties. Modeling showed that nonlinear selection on an additive functional trait of a protein is generically expected to give rise to a functional sector. These modeling results motivated a principled method, called ICOD, which is designed to identify functional sectors, as well as mutational effects, from sequence data. However, a challenge for all methods aiming to identify sectors from multiple sequence alignments is that correlations in amino-acid usage can also arise from the mere fact that homologous sequences share common ancestry, i.e. from phylogeny. Here, we generate controlled synthetic data from a minimal model comprising both phylogeny and functional sectors. We use this data to dissect the impact of phylogeny on sector identification and on mutational effect inference by different methods. We find that ICOD is most robust to phylogeny, but that conservation is also quite robust. Next, we consider natural multiple sequence alignments of protein families for which deep mutational scan experimental data is available. We show that in this natural data, conservation and ICOD best identify sites with strong functional roles, in agreement with our results on synthetic data. Importantly, these two methods have different premises, since they respectively focus on conservation and on correlations. Thus, their joint use can reveal complementary information.

q-bio.PE

Mutant fate in spatially structured populations on graphs: connecting models to experiments

In nature, most microbial populations have complex spatial structures that can affect their evolution. Evolutionary graph theory predicts that some spatial structures modelled by placing individuals on the nodes of a graph affect the probability that a mutant will fix. Evolution experiments are beginning to explicitly address the impact of graph structures on mutant fixation. However, the assumptions of evolutionary graph theory differ from the conditions of modern evolution experiments, making the comparison between theory and experiment challenging. Here, we aim to bridge this gap by using our new model of spatially structured populations. This model considers connected subpopulations that lie on the nodes of a graph, and allows asymmetric migrations. It can handle large populations, and explicitly models serial passage events with migrations, thus closely mimicking experimental conditions. We analyze recent experiments in light of this model. We suggest useful parameter regimes for future experiments, and we make quantitative predictions for these experiments. In particular, we propose experiments to directly test our recent prediction that the star graph with asymmetric migrations suppresses natural selection and can accelerate mutant fixation or extinction, compared to a well-mixed population.

q-bio.PE

Evolution of cooperation in deme-structured populations on graphs

Understanding how cooperation can evolve in populations despite its cost to individual cooperators is an important challenge. Models of spatially structured populations with one individual per node of a graph have shown that cooperation, modeled via the prisoner's dilemma, can be favored by natural selection. These results depend on microscopic update rules, which determine how birth, death and migration on the graph are coupled. Recently, we developed coarse-grained models of spatially structured populations on graphs, where each node comprises a well-mixed deme, and where migration is independent from division and death, thus bypassing the need for update rules. Here, we study the evolution of cooperation in these models in the rare migration regime, within the prisoner's dilemma. We find that cooperation is not favored by natural selection in these coarse-grained models on graphs where overall deme fitness does not directly impact migration from a deme. This is due to a separation of scales, whereby cooperation occurs at a local level within demes, while spatial structure matters between demes.

q-bio.PE

Pairing interacting protein sequences using masked language modeling

Predicting which proteins interact together from amino-acid sequences is an important task. We develop a method to pair interacting protein sequences which leverages the power of protein language models trained on multiple sequence alignments, such as MSA Transformer and the EvoFormer module of AlphaFold. We formulate the problem of pairing interacting partners among the paralogs of two protein families in a differentiable way. We introduce a method called DiffPALM that solves it by exploiting the ability of MSA Transformer to fill in masked amino acids in multiple sequence alignments using the surrounding context. MSA Transformer encodes coevolution between functionally or structurally coupled amino acids. We show that it captures inter-chain coevolution, while it was trained on single-chain data, which means that it can be used out-of-distribution. Relying on MSA Transformer without fine-tuning, DiffPALM outperforms existing coevolution-based pairing methods on difficult benchmarks of shallow multiple sequence alignments extracted from ubiquitous prokaryotic protein datasets. It also outperforms an alternative method based on a state-of-the-art protein language model trained on single sequences. Paired alignments of interacting protein sequences are a crucial ingredient of supervised deep learning methods to predict the three-dimensional structure of protein complexes. DiffPALM substantially improves the structure prediction of some eukaryotic protein complexes by AlphaFold-Multimer, without significantly deteriorating any of those we tested. It also achieves competitive performance with using orthology-based pairing.

q-bio.BM

Universal Casimir attraction between filaments at the cell scale

The electromagnetic Casimir interaction between dielectric objects immersed in salted water includes a universal contribution that is not screened by the solvent and therefore long-ranged. Here, we study the geometry of two parallel dielectric cylinders. We derive the Casimir free energy by using the scattering method. We show that its magnitude largely exceeds the thermal energy scale for a large parameter range. This includes length scales relevant for actin filaments and microtubules in cells. We show that the Casimir free energy is a universal function of the geometry, independent of the dielectric response functions of the cylinders, at all distances of biological interest. While multiple interactions exist between filaments in cells, this universal attractive interaction should have an important role in the cohesion of bundles of parallel filaments.

physics.bio-ph

Frequent asymmetric migrations suppress natural selection in spatially structured populations

Natural microbial populations often have complex spatial structures. This can impact their evolution, in particular the ability of mutants to take over. While mutant fixation probabilities are known to be unaffected by sufficiently symmetric structures, evolutionary graph theory has shown that some graphs can amplify or suppress natural selection, in a way that depends on microscopic update rules. We propose a model of spatially structured populations on graphs directly inspired by batch culture experiments, alternating within-deme growth on nodes and migration-dilution steps, and yielding successive bottlenecks. This setting bridges models from evolutionary graph theory with Wright-Fisher models. Using a branching process approach, we show that spatial structure with frequent migrations can only yield suppression of natural selection. More precisely, in this regime, circulation graphs, where the total incoming migration flow equals the total outgoing one in each deme, do not impact fixation probability, while all other graphs strictly suppress selection. Suppression becomes stronger as the asymmetry between incoming and outgoing migrations grows. Amplification of natural selection can nevertheless exist in a restricted regime of rare migrations and very small fitness advantages, where we recover the predictions of evolutionary graph theory for the star graph.

q-bio.PE

A universal attractive interaction between filaments at the cell scale

Actin filaments and microtubules both often form bundles of parallel filaments within cells. Here, we shed light on a universal attractive interaction between two such parallel filaments. Indeed, the electrodynamic Casimir interaction between dielectric objects immersed in salted water at room or body temperature includes a universal contribution that is unscreened by the solvent and therefore long-ranged. We study this interaction between two parallel cylinders immersed in salted water with strong Debye screening. We show that its magnitude can largely exceed the energy scale of thermal fluctuations in the case of actin filaments and microtubules in cells. While multiple interactions exist between filaments in cells, this universal attractive interaction should thus have an important role, e.g. in bundle formation and cohesion.

physics.bio-ph

Impact of phylogeny on structural contact inference from protein sequence data

Local and global inference methods have been developed to infer structural contacts from multiple sequence alignments of homologous proteins. They rely on correlations in amino-acid usage at contacting sites. Because homologous proteins share a common ancestry, their sequences also feature phylogenetic correlations, which can impair contact inference. We investigate this effect by generating controlled synthetic data from a minimal model where the importance of contacts and of phylogeny can be tuned. We demonstrate that global inference methods, specifically Potts models, are more resilient to phylogenetic correlations than local methods, based on covariance or mutual information. This holds whether or not phylogenetic corrections are used, and may explain the success of global methods. We analyse the roles of selection strength and of phylogenetic relatedness. We show that sites that mutate early in the phylogeny yield false positive contacts. We consider natural data and realistic synthetic data, and our findings generalise to these cases. Our results highlight the impact of phylogeny on contact prediction from protein sequences and illustrate the interplay between the rich structure of biological data and inference.

q-bio.BM

Protein language models trained on multiple sequence alignments learn phylogenetic relationships

Self-supervised neural language models with attention have recently been applied to biological sequence data, advancing structure, function and mutational effect prediction. Some protein language models, including MSA Transformer and AlphaFold's EvoFormer, take multiple sequence alignments (MSAs) of evolutionarily related proteins as inputs. Simple combinations of MSA Transformer's row attentions have led to state-of-the-art unsupervised structural contact prediction. We demonstrate that similarly simple, and universal, combinations of MSA Transformer's column attentions strongly correlate with Hamming distances between sequences in MSAs. Therefore, MSA-based language models encode detailed phylogenetic relationships. We further show that these models can separate coevolutionary signals encoding functional and structural constraints from phylogenetic correlations reflecting historical contingency. To assess this, we generate synthetic MSAs, either without or with phylogeny, from Potts models trained on natural MSAs. We find that unsupervised contact prediction is substantially more resilient to phylogenetic noise when using MSA Transformer versus inferred Potts models.

q-bio.BM

Combining phylogeny and coevolution improves the inference of interaction partners among paralogous proteins

Predicting protein-protein interactions from sequences is an important goal of computational biology. Various sources of information can be used to this end. Starting from the sequences of two interacting protein families, one can use phylogeny or residue coevolution to infer which paralogs are specific interaction partners within each species. We show that these two signals can be combined to improve the performance of the inference of interaction partners among paralogs. For this, we first align the sequence-similarity graphs of the two families through simulated annealing, yielding a robust partial pairing. We next use this partial pairing to seed a coevolution-based iterative pairing algorithm. This combined method improves performance over either separate method. The improvement obtained is striking in the difficult cases where the average number of paralogs per species is large or where the total number of sequences is modest.

q-bio.BM