Searcharxiv⌕ Search

arXiv subjects

Eugene Shakhnovich

Publications and source records attributed to Eugene Shakhnovich.

At least 19 recordsLinked to original sources

Viral Proteins Reveal Geometry of Protein Language Models

Protein language models are trained on highly imbalanced datasets, raising the question of how they represent underrepresented biological sequences. Using viral proteins as a case study across ESM model families, we identify a dominant nativeness axis in embedding space, aligned with masked reconstruction perplexity, that orders sequences from well-modeled cellular proteins through viral proteins to shuffled and random sequences. Scaling contracts this axis unevenly across viral families. Despite this, protein language model embeddings retain viral-specific signal: viral proteins remain linearly separable beyond zero-shot perplexity and shallow sequence features. Together, these results suggest that pLM representations are structured by a general notion of nativeness while preserving information specific to distinct biological groups.

cs.LG↗

Computational Design and Experimental Validation of Photoactive PARP1 Inhibitors

Light-activated drugs are a promising way to treat localized diseases for which existing treatments have severe side effects. However, their development is complicated by the set of photophysical and biological properties that must be simultaneously optimized. Here we used computational techniques to find a set of promising candidates for the photoactive inhibition of the poly(ADP-ribose) polymerase 1 (PARP1) cancer target. Using our recently developed methods based on atomistic simulation and machine learning (ML), we screened a set of 5 million hypothetical photoactive ligands. Our workflow used protein-ligand docking to identify candidates with differential PARP1 binding under light and dark conditions; ML force fields and quantum chemistry calculations to predict p$K_\mathrm{a}$, absorption spectra, and thermal half-lives; graph-based surrogate models to screen additional compounds; excited-state nonadiabatic dynamics with ML force fields to estimate quantum yields; and free energy perturbation (FEP) to refine binding predictions. From these predictions, we prioritized a small set of synthetically feasible candidates expected to have red-shifted absorption spectra, thermal half-lives on the order of seconds to minutes, and isomer-dependent PARP1 binding under visible-light control. We synthesized 10 candidates and experimentally characterized their photobehavior and PARP1 inhibition constants. Among the validated compounds, \textbf{1} showed a 15-fold increase in inhibition of PARP1 upon green-light irradiation at 519 nm (208.8 $\pm$ 28.3 $μ$M vs 14.4 $\pm$ 1.9 $μ$M). These results validate the computation-guided screening strategy for identifying red-shifted PARP1 photoinhibitors, while also underscoring current limitations such as rapid thermal relaxation in aqueous media.

physics.chem-ph↗

Evolutionary Dynamics of a Lattice Dimer: a Toy Model for Stability vs. Affinity Trade-offs in Proteins

Understanding how a stressor applied on a biological system shapes its evolution is key to achieving targeted evolutionary control. Here we present a toy model of two interacting lattice proteins to quantify the response to the selective pressure defined by the binding energy. We generate sequence data of proteins and study how the sequence and structural properties of dimers are affected by the applied selective pressure, both during the evolutionary process and in the stationary regime. In particular we show that internal contacts of native structures lose strength, while inter-structure contacts are strengthened due to the folding-binding competition. We discuss how dimerization is achieved through enhanced mutability on the interacting faces, and how the designability of each native structure changes upon introduction of the stressor.

cond-mat.stat-mech↗

Mapping the space of photoswitchable ligands and photodruggable proteins with computational modeling

Light-activated drugs are a promising way to localize biological activity and minimize side effects. However, their development is complicated by the numerous photophysical and biological properties that must be simultaneously optimized. To accelerate the design of photoactive drugs, we describe a procedure that combines ligand-protein docking with chemical property prediction based on machine learning (ML). We apply this procedure to 58 proteins and 9,000 photo-drug candidates based on azobenzene cis-trans isomerism. We find that most proteins display a preference for trans isomers over cis, and that the binding affinities of nominally active/inactive pairs are in fact highly correlated. These findings have significant value for photopharmacology research, and reinforce the need for virtual screening to identify compounds with rare desirable properties. Further, we combine our procedure with quantum chemical validation to identify promising candidates for the photoactive inhibition of PARP1, an enzyme that is over-expressed in cancer cells. The top compounds are predicted to have long-lived active forms, differential bioactivity, and absorption in the near-infrared therapeutic window.

physics.chem-ph↗

Thermal half-lives of azobenzene derivatives: virtual screening based on intersystem crossing using a machine learning potential

Molecular photoswitches are the foundation of light-activated drugs. A key photoswitch is azobenzene, which exhibits trans-cis isomerism in response to light. The thermal half-life of the cis isomer is of crucial importance, since it controls the duration of the light-induced biological effect. Here we introduce a computational tool for predicting the thermal half-lives of azobenzene derivatives. Our automated approach uses a fast and accurate machine learning potential trained on quantum chemistry data. Building on well-established earlier evidence, we argue that thermal isomerization proceeds through rotation mediated by intersystem crossing, and incorporate this mechanism into our automated workflow. We use our approach to predict the thermal half-lives of 19,000 azobenzene derivatives. We explore trends and tradeoffs between barriers and absorption wavelengths, and open-source our data and software to accelerate research in photopharmacology.

physics.chem-ph↗

Excited state, non-adiabatic dynamics of large photoswitchable molecules using a chemically transferable machine learning potential

Light-induced chemical processes are ubiquitous in nature and have widespread technological applications. For example, photoisomerization can allow a drug with a photo-switchable scaffold such as azobenzene to be activated with light. In principle, photoswitches with desired photophysical properties like high isomerization quantum yields can be identified through virtual screening with reactive simulations. In practice, these simulations are rarely used for screening, since they require hundreds of trajectories and expensive quantum chemical methods to account for non-adiabatic excited state effects. Here we introduce a diabatic artificial neural network (DANN) based on diabatic states to accelerate such simulations for azobenzene derivatives. The network is six orders of magnitude faster than the quantum chemistry method used for training. DANN is transferable to azobenzene molecules outside the training set, predicting quantum yields for unseen species that are correlated with experiment. We use the model to virtually screen 3,100 hypothetical molecules, and identify novel species with extremely high predicted quantum yields. The model predictions are confirmed using high accuracy non-adiabatic dynamics. Our results pave the way for fast and accurate virtual screening of photoactive compounds.

physics.chem-ph↗

Development of antibacterial compounds that block evolutionary pathways to resistance

Antibiotic resistance is a worldwide challenge. A potential approach to block resistance is to simultaneously inhibit WT and known escape variants of the target bacterial protein. Here we applied an integrated computational and experimental approach to discover compounds that inhibit both WT and trimethoprim (TMP) resistant mutants of E. coli dihydrofolate reductase (DHFR). We identified a novel compound (CD15-3) that inhibits WT DHFR and its TMP resistant variants L28R, P21L and A26T with IC50 50-75 micromoles against WT and TMP-resistant strains. Resistance to CD15-3 was dramatically delayed compared to TMP in in vitro evolution. Whole genome sequencing of CD15-3 resistant strains showed no mutations in the target folA locus. Rather, gene duplication of several efflux pumps gave rise to weak (about twofold increase in IC50) resistance against CD15-3. Altogether, our results demonstrate the promise of strategy to develop evolution drugs - compounds which block evolutionary escape routes in pathogens.

q-bio.BM↗

Liquid-liquid microphase separation leads to formation of membraneless organelles

Proteins and nucleic acids can spontaneously self-assemble into membraneless droplet-like compartments, both in vitro and in vivo. A key component of these droplets are multi-valent proteins that possess several adhesive domains with specific interaction partners (whose number determines total valency of the protein) separated by disordered regions. Here, using multi-scale simulations we show that such proteins self-organize into micro-phase separated droplets of various sizes as opposed to the Flory-like macro-phase separated equilibrium state of homopolymers or equilibrium physical gels. We show that the micro-phase separated state is a dynamic outcome of the interplay between two competing processes: a diffusion-limited encounter between proteins, and the dynamics within small clusters that results in exhaustion of available valencies whereby all specifically interacting domains find their interacting partners within smaller clusters, leading to arrested phase separation. We first model these multi-valent chains as bead-spring polymers with multiple adhesive domains separated by semi-flexible linkers and use Langevin Dynamics (LD) to assess how key timescales depend on the molecular properties of associating polymers. Using the time-scales from LD simulations, we develop a coarse-grained kinetic model to study this phenomenon at longer times. Consistent with LD simulations, the macro-phase separated state was only observed at high concentrations and large interaction valencies. Further, in the regime where cluster sizes approach macro-phase separation, the condensed phase becomes dynamically solid-like, suggesting that it might no longer be biologically functional. Therefore, the micro-phase separated state could be a hallmark of functional droplets formed by proteins with the sticker-spacer architecture.

q-bio.BM↗

Evolutionary dynamics determines adaptation to inactivation of an essential gene

Genetic inactivation of essential genes creates an evolutionary scenario distinct from escape from drug inhibition, but the mechanisms of microbe adaptations in such cases remain unknown. Here we inactivate E. coli dihydrofolate reductase (DHFR) by introducing D27G,N,F chromosomal mutations in a key catalytic residue with subsequent adaptation by serial dilutions. The partial reversal G27->C occurred in three evolutionary trajectories. Conversely, in one trajectory for D27G and in all trajectories for D27F,N strains adapted to grow at very low supplement folAmix concentrations but did not escape entirely from supplement auxotrophy. Major global shifts in metabolome and proteome occurred upon DHFR inactivation, which were partially reversed in adapted strains. Loss of function mutations in two genes, thyA and deoB, ensured adaptation to low folAmix by rerouting the 2-Deoxy-D-ribose-phosphate metabolism from glycolysis towards synthesis of dTMP. Multiple evolutionary pathways of adaptation to low folAmix converge to highly accessible yet suboptimal fitness peak.

q-bio.PE↗

Statistical Physics of the Symmetric Group

Ordered chains (such as chains of amino acids) are ubiquitous in biological cells, and these chains perform specific functions contingent on the sequence of their components. Using the existence and general properties of such sequences as a theoretical motivation, we study the statistical physics of systems whose state space is defined by the possible permutations of an ordered list, i.e., the symmetric group, and whose energy is a function of how certain permutations deviate from some chosen correct ordering. Such a non-factorizable state space is quite different from the state spaces typically considered in statistical physics systems and consequently has novel behavior in systems with interacting and even non-interacting Hamiltonians. Various parameter choices of a mean-field model reveal the system to contain five different physical regimes defined by two transition temperatures, a triple point, and a quadruple point. Finally, we conclude by discussing how the general analysis can be extended to state spaces with more complex combinatorial properties and to other standard questions of statistical mechanics models.

cond-mat.stat-mech↗

Benchmarking inverse statistical approaches for protein structure and design with exactly solvable models

Inverse statistical approaches to determine protein structure and function from Multiple Sequence Alignments (MSA) are emerging as powerful tools in computational biology. However the underlying assumptions of the relationship between the inferred effective Potts Hamiltonian and real protein structure and energetics remain untested so far. Here we use lattice protein model (LP) to benchmark those inverse statistical approaches. We build MSA of highly stable sequences in target LP structures, and infer the effective pairwise Potts Hamiltonians from those MSA. We find that inferred Potts Hamiltonians reproduce many important aspects of 'true' LP structures and energetics. Careful analysis reveals that effective pairwise couplings in inferred Potts Hamiltonians depend not only on the energetics of the native structure but also on competing folds; in particular, the coupling values reflect both positive design (stabilization of native conformation) and negative design (destabilization of competing folds). In addition to providing detailed structural information, the inferred Potts models used as protein Hamiltonian for design of new sequences are able to generate with high probability completely new sequences with the desired folds, which is not possible using independent-site models. Those are remarkable results as the effective LP Hamiltonians used to generate MSA are not simple pairwise models due to the competition between the folds. Our findings elucidate the reasons for the success of inverse approaches to the modelling of proteins from sequence data, and their limitations.

q-bio.BM↗

Interplay between pleiotropy and secondary selection determines rise and fall of mutators in stress response

Dramatic rise of mutators has been found to accompany adaptation of bacteria in response to many kinds of stress. Two views on the evolutionary origin of this phenomenon emerged: the pleiotropic hypothesis positing that it is a byproduct of environmental stress or other specific stress response mechanisms and the second order selection which states that mutators hitchhike to fixation with unrelated beneficial alleles. Conventional population genetics models could not fully resolve this controversy because they are based on certain assumptions about fitness landscape. Here we address this problem using a microscopic multiscale model, which couples physically realistic molecular descriptions of proteins and their interactions with population genetics of carrier organisms without assuming any a priori fitness landscape. We found that both pleiotropy and second order selection play a crucial role at different stages of adaptation: the supply of mutators is provided through destabilization of error correction complexes or fluctuations of production levels of prototypic mismatch repair proteins (pleiotropic effects), while rise and fixation of mutators occur when there is a sufficient supply of beneficial mutations in replication-controlling genes. This general mechanism assures a robust and reliable adaptation of organisms to unforeseen challenges. This study highlights physical principles underlying physical biological mechanisms of stress response and adaptation.

q-bio.BM↗

Adaptation through stochastic switching into transient mutators in finite asexual populations

The importance of mutator clones in the adaptive evolution of asexual populations is not fully understood. Here we address this problem by using an ab initio microscopic model of living cells, whose fitness is derived directly from their genomes using a biophysically realistic model of protein folding and interactions in the cytoplasm. The model organisms contain replication controlling genes (DCGs) and genes modeling the mismatch repair (MMR) complexes. We find that adaptation occurs through the transient fixation of a mutator phenotype, regardless of particular perturbations in the fitness landscape. The microscopic pathway of adaptation follows a well-defined set of events: stochastic switching to the mutator phenotype first, then mutation in the MMR complex that hitchhikes with a beneficial mutation in the DCGs, and finally a compensating mutation in the MMR complex returning the population to a non-mutator phenotype. Similarity of these results to reported adaptation events points out to robust universal physical principles of evolutionary adaptation.

q-bio.BM↗

Universality and diversity of folding mechanics for three-helix bundle proteins

In this study we evaluate, at full atomic detail, the folding processes of two small helical proteins, the B domain of protein A and the Villin headpiece. Folding kinetics are studied by performing a large number of ab initio Monte Carlo folding simulations using a single transferable all-atom potential. Using these trajectories, we examine the relaxation behavior, secondary structure formation, and transition-state ensembles (TSEs) of the two proteins and compare our results with experimental data and previous computational studies. To obtain a detailed structural information on the folding dynamics viewed as an ensemble process, we perform a clustering analysis procedure based on graph theory. Moreover, rigorous pfold analysis is used to obtain representative samples of the TSEs and a good quantitative agreement between experimental and simulated Fi-values is obtained for protein A. Fi-values for Villin are also obtained and left as predictions to be tested by future experiments. Our analysis shows that two-helix hairpin is a common partially stable structural motif that gets formed prior to entering the TSE in the studied proteins. These results together with our earlier study of Engrailed Homeodomain and recent experimental studies provide a comprehensive, atomic-level picture of folding mechanics of three-helix bundle proteins.

q-bio.BM↗

The Hypercube of Life: How Protein Stability Imposes Limits on Organism Complexity and Speed of Molecular Evolution

Classical population genetics a priori assigns fitness to alleles without considering molecular or functional properties of proteins that these alleles encode. Here we study population dynamics in a model where fitness can be inferred from physical properties of proteins under a physiological assumption that loss of stability of any protein encoded by an essential gene confers a lethal phenotype. Accumulation of mutations in organisms containing Gamma genes can then be represented as diffusion within the Gamma dimensional hypercube with adsorbing boundaries which are determined, in each dimension, by loss of a protein stability and, at higher stability, by lack of protein sequences. Solving the diffusion equation whose parameters are derived from the data on point mutations in proteins, we determine a universal distribution of protein stabilities, in agreement with existing data. The theory provides a fundamental relation between mutation rate, maximal genome size and thermodynamic response of proteins to point mutations. It establishes a universal speed limit on rate of molecular evolution by predicting that populations go extinct (via lethal mutagenesis) when mutation rate exceeds approximately 6 mutations per essential part of genome per replication for mesophilic organisms and 1 to 2 mutations per genome per replication for thermophilic ones. Further, our results suggest that in absence of error correction, modern RNA viruses and primordial genomes must necessarily be very short. Several RNA viruses function close to the evolutionary speed limit while error correction mechanisms used by DNA viruses and non-mutant strains of bacteria featuring various genome lengths and mutation rates have brought these organisms universally about 1000 fold below the natural speed limit.

q-bio.BM↗

A first-principles model of early evolution: Emergence of gene families, species and preferred protein folds

In this work we develop a microscopic physical model of early evolution, where phenotype,organism life expectancy, is directly related to genotype, the stability of its proteins in their native conformations which can be determined exactly in the model. Simulating the model on a computer, we consistently observe the Big Bang scenario whereby exponential population growth ensues as soon as favorable sequence-structure combinations (precursors of stable proteins) are discovered. Upon that, random diversity of the structural space abruptly collapses into a small set of preferred proteins. We observe that protein folds remain stable and abundant in the population at time scales much greater than mutation or organism lifetime, and the distribution of the lifetimes of dominant folds in a population approximately follows a power law. The separation of evolutionary time scales between discovery of new folds and generation of new sequences gives rise to emergence of protein families and superfamilies whose sizes are power-law distributed, closely matching the same distributions for real proteins. On the population level we observe emergence of species, subpopulations which carry similar genomes. Further we present a simple theory that relates stability of evolving proteins to the sizes of emerging genomes. Together, these results provide a microscopic first principles picture of how first gene families developed in the course of early evolution

q-bio.BM↗

Robust protein-protein interactions in crowded cellular environments

The capacity of proteins to interact specifically with one another underlies our conceptual understanding of how living systems function. Systems-level study of specificity in protein-protein interactions is complicated by the fact that the cellular environment is crowded and heterogeneous; interaction pairs may exist at low relative concentrations and thus be presented with many more opportunities for promiscuous interactions compared to specific interaction possibilities. Here we address these questions using a simple computational model that includes specifically designed interacting model proteins immersed in a mixture containing hundreds of different unrelated ones; all of them undergo simulated diffusion and interaction. We find that specific complexes are quite robust to interference from promiscuous interaction partners, only in the range of temperatures Tdesign>T>Trand. At T>Tdesign specific complexes become unstable, while at T Trand. This condition requires an energy gap between binding energy in a specific complex and set of binding energies between randomly associating proteins, providing a general physical constraint on evolutionary selection or design of specific interacting protein interfaces. This work has implications for our understanding of how the protein repertoire functions and evolves within the context of cellular systems.

q-bio.BM↗

Protein Structure and Evolutionary History Determine Sequence Space Topology

Understanding the observed variability in the number of homologs of a gene is a very important, unsolved problem that has broad implications for research into co-evolution of structure and function, gene duplication, pseudogene formation and possibly for emerging diseases. Here we attempt to define and elucidate the reasons behind this observed unevenness in sequence space. We present evidence that sequence variability and functional diversity of a gene or fold family is influenced by certain quantitative characteristics of the protein structure that reflect potential for sequence plasticity i.e. the ability to accept mutation without losing thermodynamic stability.

q-bio.BM↗