SearcharxivSearch

arXiv subjects

Jakub Otwinowski

Publications and source records attributed to Jakub Otwinowski.

15 recordsLinked to original sources

Contrastive losses as generalized models of global epistasis

Fitness functions map large combinatorial spaces of biological sequences to properties of interest. Inferring these multimodal functions from experimental data is a central task in modern protein engineering. Global epistasis models are an effective and physically-grounded class of models for estimating fitness functions from observed data. These models assume that a sparse latent function is transformed by a monotonic nonlinearity to emit measurable fitness. Here we demonstrate that minimizing supervised contrastive loss functions, such as the Bradley-Terry loss, is a simple and flexible technique for extracting the sparse latent function implied by global epistasis. We argue by way of a fitness-epistasis uncertainty principle that the nonlinearities in global epistasis models can produce observed fitness functions that do not admit sparse representations, and thus may be inefficient to learn from observations when using a Mean Squared Error (MSE) loss (a common practice). We show that contrastive losses are able to accurately estimate a ranking function from limited data even in regimes where MSE is ineffective and validate the practical utility of this insight by demonstrating that contrastive loss functions result in consistently improved performance on benchmark tasks.

q-bio.PE

Learning the shape of protein micro-environments with a holographic convolutional neural network

Proteins play a central role in biology from immune recognition to brain activity. While major advances in machine learning have improved our ability to predict protein structure from sequence, determining protein function from structure remains a major challenge. Here, we introduce Holographic Convolutional Neural Network (H-CNN) for proteins, which is a physically motivated machine learning approach to model amino acid preferences in protein structures. H-CNN reflects physical interactions in a protein structure and recapitulates the functional information stored in evolutionary data. H-CNN accurately predicts the impact of mutations on protein function, including stability and binding of protein complexes. Our interpretable computational model for protein structure-function maps could guide design of novel proteins with desired function.

physics.bio-ph

Design of an optimal combination therapy with broadly neutralizing antibodies to suppress HIV-1

Broadly neutralizing antibodies (bNAbs) are promising targets for vaccination and therapy against HIV. Passive infusions of bNAbs have shown promise in clinical trials as a potential alternative for anti-retroviral therapy. A key challenge for the potential clinical application of bnAbs is the suppression of viral escape, which is more effectively achieved with a combination of bNAbs. However, identifying an optimal bNAb cocktail is combinatorially complex. Here, we propose a computational approach to predict the efficacy of a bNAb therapy trial based on the population genetics of HIV escape, which we parametrize using high-throughput HIV sequence data from a cohort of untreated bNAb-naive patients. By quantifying the mutational target size and the fitness cost of HIV-1 escape from bNAbs, we reliably predict the distribution of rebound times in three clinical trials. Importantly, we show that early rebounds are dominated by the pre-treatment standing variation of HIV-1 populations, rather than spontaneous mutations during treatment. Lastly, we show that a cocktail of three bNAbs is necessary to suppress the chances of viral escape below 1%, and we predict the optimal composition of such a bNAb cocktail. Our results offer a rational design for bNAb therapy against HIV-1, and more generally show how genetic data could be used to predict treatment outcomes and design new approaches to pathogenic control.

q-bio.PE

Dynamics of B-cell repertoires and emergence of cross-reactive responses in COVID-19 patients with different disease severity

COVID-19 patients show varying severity of the disease ranging from asymptomatic to requiring intensive care. Although a number of SARS-CoV-2 specific monoclonal antibodies have been identified, we still lack an understanding of the overall landscape of B-cell receptor (BCR) repertoires in COVID-19 patients. Here, we used high-throughput sequencing of bulk and plasma B-cells collected over multiple time points during infection to characterize signatures of B-cell response to SARS-CoV-2 in 19 patients. Using principled statistical approaches, we determined differential features of BCRs associated with different disease severity. We identified 38 significantly expanded clonal lineages shared among patients as candidates for specific responses to SARS-CoV-2. Using single-cell sequencing, we verified reactivity of BCRs shared among individuals to SARS-CoV-2 epitopes. Moreover, we identified natural emergence of a BCR with cross-reactivity to SARS-CoV-1 and SARS-CoV-2 in a number of patients. Our results provide important insights for development of rational therapies and vaccines against COVID-19.

q-bio.GN

Fierce selection and interference in B-cell repertoire response to chronic HIV-1

During chronic infection, HIV-1 engages in a rapid coevolutionary arms race with the host's adaptive immune system. While it is clear that HIV exerts strong selection on the adaptive immune system, the characteristics of the somatic evolution that shape the immune response are still unknown. Traditional population genetics methods fail to distinguish chronic immune response from healthy repertoire evolution. Here, we infer the evolutionary modes of B-cell repertoires and identify complex dynamics with a constant production of better B-cell receptor mutants that compete, maintaining large clonal diversity and potentially slowing down adaptation. A substantial fraction of mutations that rise to high frequencies in pathogen engaging CDRs of B-cell receptors (BCRs) are beneficial, in contrast to many such changes in structurally relevant frameworks that are deleterious and circulate by hitchhiking. We identify a pattern where BCRs in patients who experience larger viral expansions undergo stronger selection with a rapid turnover of beneficial mutations due to clonal interference in their CDR3 regions. Using population genetics modeling, we show that the extinction of these beneficial mutations can be attributed to the rise of competing beneficial alleles and clonal interference. The picture is of a dynamic repertoire, where better clones may be outcompeted by new mutants before they fix.

q-bio.PE

Information-geometric optimization with natural selection

Evolutionary algorithms, inspired by natural evolution, aim to optimize difficult objective functions without computing derivatives. Here we detail the relationship between population genetics and evolutionary optimization and formulate a new evolutionary algorithm. Optimization of a continuous objective function is analogous to searching for high fitness phenotypes on a fitness landscape. We summarize how natural selection moves a population along the non-euclidean gradient that is induced by the population on the fitness landscape (the natural gradient). Under normal approximations common in quantitative genetics, we show how selection is related to Newton's method in optimization. We find that intermediate selection is most informative of the fitness landscape. We describe the generation of new phenotypes and introduce an operator that recombines the whole population to generate variants that preserve normal statistics. Finally, we introduce a proof-of-principle algorithm that combines natural selection, our recombination operator, and an adaptive method to increase selection. Our algorithm is similar to covariance matrix adaptation and natural evolutionary strategies in optimization, and has similar performance. The algorithm is extremely simple in implementation with no matrix inversion or factorization, does not require storing a covariance matrix, and may form the basis of more general model-based optimization algorithms with natural gradient updates.

q-bio.PE

Biophysical inference of epistasis and the effects of mutations on protein stability and function

Understanding the relationship between protein sequence, function, and stability is a fundamental problem in biology. While high-throughput methods have produced large numbers of sequence-function pairs, functional assays do not distinguish whether mutations directly affect function or are destabilizing the protein. Here, we introduce a statistical method to infer the underlying biophysics from a high-throughput binding assay by combining information from many mutated variants. We fit a thermodynamic model describing the bound, unbound, and unfolded states to high quality data of protein G domain B1 binding to IgG-Fc. We infer an energy landscape with distinct folding and binding energies for each substitution providing a detailed view of how mutations affect binding and stability across the protein. We accurately infer folding energy of each variant in physical units, validated by independent data, whereas previous high-throughput methods could only measure indirect changes in stability. While we assume an additive sequence-energy relationship, the binding fraction is epistatic due its non-linear relation to energy. Despite having no epistasis in energy, our model explains much of the observed epistasis in binding fraction, with the remaining epistasis identifying conformationally dynamic regions.

q-bio.BM

Host-pathogen coevolution and the emergence of broadly neutralizing antibodies in chronic infections

The vertebrate adaptive immune system provides a flexible and diverse set of molecules to neutralize pathogens. Yet, viruses such as HIV can cause chronic infections by evolving as quickly as the adaptive immune system, forming an evolutionary arms race. Here we introduce a mathematical framework to study the coevolutionary dynamics of antibodies with antigens within a host. We focus on changes in the binding interactions between the antibody and antigen populations, which result from the underlying stochastic evolution of genotype frequencies driven by mutation, selection, and drift. We identify the critical viral and immune parameters that determine the distribution of antibody-antigen binding affinities. We also identify definitive signatures of coevolution that measure the reciprocal response between antibodies and viruses, and we introduce experimentally measurable quantities that quantify the extent of adaptation during continual coevolution of the two opposing populations. Using this analytical framework, we infer rates of viral and immune adaptation based on time-shifted neutralization assays in two HIV-infected patients. Finally, we analyze competition between clonal lineages of antibodies and characterize the fate of a given lineage in terms of the state of the antibody and viral populations. In particular, we derive the conditions that favor the emergence of broadly neutralizing antibodies, which may be useful in designing a vaccine against HIV.

q-bio.PE

The diversity of evolutionary dynamics on epistatic versus non-epistatic fitness landscapes

The class of epistatic fitness landscapes is much more diverse than the class of non-epistatic landscapes, and so it stands to reason that there exist dynamical phenomena that can only be realized in the presence of epistasis. Here, we compare evolutionary dynamics on all finite epistatic landscapes versus all finite non-epistatic landscapes, under weak mutation. We first analyze the mean fitness trajectory - that is, the time course of the expected fitness of a population. We show that for any epistatic fitness landscape and starting genotype, there always exists a non-epistatic fitness landscape and starting genotype that produces the exact same mean fitness trajectory. Thus, surprisingly, the space of mean fitness trajectories that can be realized by epistatic landscapes is no more diverse than the space of mean fitness trajectories that can be realized by non-epistatic landscapes. On the other hand, we show that epistatic fitness landscapes can produce dynamics in the time-evolution of the variance in fitness across replicate populations and in the time-evolution of the expected number of substitutions that cannot be produced by any non-epistatic landscape. These results on identifiability have implications for efforts to infer epistasis from the types of data often measured in experimental populations.

q-bio.PE

Clonal interference and Muller's ratchet in spatial habitats

Competition between independently arising beneficial mutations is enhanced in spatial populations due to the linear rather than exponential growth of clones. Recent theoretical studies have pointed out that the resulting fitness dynamics is analogous to a surface growth process, where new layers nucleate and spread stochastically, leading to the build up of scale-invariant roughness. This scenario differs qualitatively from the standard view of adaptation in that the speed of adaptation becomes independent of population size while the fitness variance does not. Here we exploit recent progress in the understanding of surface growth processes to obtain precise predictions for the universal, non-Gaussian shape of the fitness distribution for one-dimensional habitats, which are verified by simulations. When the mutations are deleterious rather than beneficial the problem becomes a spatial version of Muller's ratchet. In contrast to the case of well-mixed populations, the rate of fitness decline remains finite even in the limit of an infinite habitat, provided the ratio $U_d/s^2$ between the deleterious mutation rate and the square of the (negative) selection coefficient is sufficiently large. Using again an analogy to surface growth models we show that the transition between the stationary and the moving state of the ratchet is governed by directed percolation.

q-bio.PE

Inferring fitness landscapes by regression produces biased estimates of epistasis

The genotype-fitness map plays a fundamental role in shaping the dynamics of evolution. However, it is difficult to directly measure a fitness landscape in practice, because the number of possible genotypes is astronomical. One approach is to sample as many genotypes as possible, measure their fitnesses, and fit a statistical model of the landscape that includes additive and pairwise interactive effects between loci. Here we elucidate the pitfalls of using such regressions, by studying artificial but mathematically convenient fitness landscapes. We identify two sources of bias inherent in these regression procedures that each tends to under-estimate high fitnesses and over-estimate low fitnesses. We characterize these biases for random sampling of genotypes, as well as for samples drawn from a population under selection in the Wright-Fisher model of evolutionary dynamics. We show that common measures of epistasis, such as the number of monotonically increasing paths between ancestral and derived genotypes, the prevalence of sign epistasis, and the number of local fitness maxima, are distorted in the inferred landscape. As a result, the inferred landscape will provide systematically biased predictions for the dynamics of adaptation. We identify the same biases in a computational RNA-folding landscape, as well as in regulatory sequence binding data, treated with the same fitting procedure. Finally, we present a method that may ameliorate these biases in some cases.

q-bio.PE

Accumulation of beneficial mutations in one dimension

When beneficial mutations are relatively common, competition between multiple unfixed mutations can reduce the rate of fixation in well-mixed asexual populations. We introduce a one dimensional model with a steady accumulation of beneficial mutations. We find a transition between periodic selection and multiple-mutation regimes. In the multiple-mutation regime, the increase of fitness along the lattice bears a striking similarity to surface growth phenomena, with power law growth and saturation of the interface width. We also find significant differences compared to the well-mixed model. In our lattice model, the transition between regimes happens at a much lower mutation rate due to slower fixation times in one dimension. Also the rate of fixation is reduced with increasing mutation rate due to the more intense competition, and it saturates with large population size.

q-bio.PE

Genotype to phenotype mapping and the fitness landscape of the E. coli lac promoter

Genotype-to-phenotype maps and the related fitness landscapes that include epistatic interactions are difficult to measure because of their high dimensional structure. Here we construct such a map using the recently collected corpora of high-throughput sequence data from the 75 base pairs long mutagenized E. coli lac promoter region, where each sequence is associated with its phenotype, the induced transcriptional activity measured by a fluorescent reporter. We find that the additive (non-epistatic) contributions of individual mutations account for about two-thirds of the explainable phenotype variance, while pairwise epistasis explains about 7% of the variance for the full mutagenized sequence and about 15% for the subsequence associated with protein binding sites. Surprisingly, there is no evidence for third order epistatic contributions, and our inferred fitness landscape is essentially single peaked, with a small amount of antagonistic epistasis. There is a significant selective pressure on the wild type, which we deduce to be multi-objective optimal for gene expression in environments with different nutrient sources. We identify transcription factor (CRP) and RNA polymerase binding sites in the promotor region and their interactions without difficult optimization steps. In particular, we observe evidence for previously unexplored genetic regulatory mechanisms, possibly kinetic in nature. We conclude with a cautionary note that inferred properties of fitness landscapes may be severely influenced by biases in the sequence data.

q-bio.PE

Speeding up evolutionary search by small fitness fluctuations

We consider a fixed size population that undergoes an evolutionary adaptation in the weak mutuation rate limit, which we model as a biased Langevin process in the genotype space. We show analytically and numerically that, if the fitness landscape has a small highly epistatic (rough) and time-varying component, then the population genotype exhibits a high effective diffusion in the genotype space and is able to escape local fitness minima with a large probability. We argue that our principal finding that even very small time-dependent fluctuations of fitness can substantially speed up evolution is valid for a wide class of models.

q-bio.PE

Totally Asymmetric Exclusion Process with Hierarchical Long-Range Connections

A non-equilibrium particle transport model, the totally asymmetric exclusion process, is studied on a one-dimensional lattice with a hierarchy of fixed long-range connections. This model breaks the particle-hole symmetry observed on an ordinary one-dimensional lattice and results in a surprisingly simple phase diagram, without a maximum-current phase. Numerical simulations of the model with open boundary conditions reveal a number of dynamic features and suggest possible applications.

cond-mat.stat-mech