SearcharxivSearch

arXiv subjects

Rama Ranganathan

Publications and source records attributed to Rama Ranganathan.

7 recordsLinked to original sources

Expanding functional protein sequence space using high entropy generative models

Boltzmann Machines trained on evolutionary sequence data have emerged as a powerful paradigm for the data-driven design of artificial proteins. However, the relationship between model architecture, specifically parameter density, and experimental performance remains poorly understood. Here, we investigate this relationship using the Chorismate Mutase enzyme family as a model system. We compare standard fully connected Boltzmann Machines for Direct Coupling Analysis (bmDCA) with sparse models generated via progressive edge activation (eaDCA) and edge decimation (edDCA). We identify a maximum-entropy model (meDCA) along the decimation trajectory that represents an optimal balance between constraint satisfaction and the flexibility of the probability distribution. We synthesized and tested artificial sequences from all models using an in vivo complementation assay, finding that all architectures, regardless of sparsity, generate functional enzymes with high success rates, even at significant divergence from natural sequences. Despite this functional equivalence, we demonstrate that the meDCA model samples a viable sequence space that is more than fifteen orders of magnitude larger than its low-entropy counterparts. Furthermore, comparative analyses reveal that high-entropy models systematically minimize overfitting and better capture the local neutral spaces surrounding natural proteins. These findings suggest that while various models satisfying coevolutionary statistics can generate functional sequences, high-entropy Boltzmann Machines provide a superior representation of the underlying evolutionary fitness landscape.

q-bio.QM

Simple biological controllers drive the evolution of soft modes

Biological systems, with many interacting components, face high-dimensional environmental fluctuations, ranging from diverse nutrient deprivations to toxins, drugs, and physical stresses. Yet, many biological control mechanisms are `simple' -- they restore homeostasis through low-dimensional representations of the system's high-dimensional state. How do low-dimensional controllers maintain homeostasis in high-dimensional systems? We develop an analytically tractable model of integral feedback for complex systems in fluctuating environments. We find that selection for homeostasis leads to the emergence of a soft mode that provides the dimensionality reduction required for the functioning of simple controllers. Our theory predicts that simple controllers that buffer environmental perturbations (e.g., stress response pathways) will also buffer mutational perturbation, an equivalence we test using experimental data across ~5000 strains in the yeast knockout collection. We also predict, counterintuitively, that knocking out a simple controller will \emph{decrease} the dimensionality of the response to environmental change; we outline transcriptomics tests to validate this. Our work suggests an evolutionary origin of soft modes whose function is for dimensionality reduction in and of itself rather than direct function like allostery, with implications ranging from cryptic genetic variation to global epistasis.

physics.bio-ph

A Unified View of Allostery

Allostery is a fundamental property of proteins that represents the functional coupling between distantly located sites. In different manifestations, this property underlies signal transduction, gene expression, and regulation -- elementary reactions in networks comprising cellular information and metabolic processing systems. In this work, we present a reduced theoretical model for allostery that encompasses the many diverse mechanisms described in past work. What emerges is a basic classification of different forms of allostery into two groups: those that couple distant sites through soft normal modes within a single state, and those that possess multiple states, with effective interactions between sites emerging only upon transitioning between these states. This work serves to unify the extensive past theoretical work on the phenomenology and machinery of allostery in proteins.

q-bio.BM

On the emergence of single versus multi-state allostery

Several physical mechanisms have been proposed to explain allostery in proteins. They differ by the number of internal states that they assume a protein to occupy, leaving open the question of what controls the emergence of these distinct physical forms of allostery. Here, we analyze a simplified model of protein allostery under a range of physical and evolutionary constraints. We find that two archetypal mechanisms can emerge through evolution: a single-state mechanism where ligand binding induces a displacement along a soft normal mode or a multi-state mechanism where ligand binding induces a switch across an energy barrier to a different stable state. Importantly, whenever the two mechanisms are possible, the multi-state mechanism confers a stronger allosteric effect and thus a selective advantage. This work defines the essential constraints on single or multi-state allostery, and sets the stage for a physical theory of its evolutionary origins.

q-bio.BM

The relation between alignment covariance and background-averaged epistasis

Epistasis, or the context-dependence of the effects of mutations, limits our ability to predict the functional impact of combinations of mutations, and ultimately our ability to predict evolutionary trajectories. Information about the context-dependence of mutations can essentially be obtained in two ways: First, by experimental measurement the functional effects of combinations of mutations and calculating the epistatic contributions directly, and second, by statistical analysis of the frequencies and co-occurrences of protein residues in a multiple sequence alignment of protein homologs. In this manuscript, we derive the mathematical relationship between epistasis calculated on the basis of functional measurements, and the covariance calculated from a multiple sequence alignment. There is no one-to-one mapping between covariance and epistatic terms: covariance implies epistasis, but epistasis does not necessarily lead to covariance, indicating that covariance in itself is not the directly relevant quantity for functional prediction. Having calculated epistatic contributions from the alignment, we can directly obtain a functional prediction from the alignment statistics by applying a Walsh-Hadamard transform, fully analogous to the transformation that reconstructs functional data from measured epistatic contributions. This embedding into the Hadamard framework is directly relevant for solidifying our theoretical understanding of statistical methods that predict function and three-dimensional structure from natural alignments.

q-bio.QM

The context-dependence of mutations: a linkage of formalisms

Defining the extent of epistasis - the non-independence of the effects of mutations - is essential for understanding the relationship of genotype, phenotype, and fitness in biological systems. The applications cover many areas of biological research, including biochemistry, genomics, protein and systems engineering, medicine, and evolutionary biology. However, the quantitative definitions of epistasis vary among fields, and its analysis beyond just pairwise effects remains obscure in general. Here, we show that different definitions of epistasis are versions of a single mathematical formalism - the weighted Walsh-Hadamard transform. We discuss that one of the definitions, the backgound-averaged epistasis, is the most informative when the goal is to uncover the general epistatic structure of a biological system, a description that can be rather different from the local epistatic structure of specific model systems. Key issues are the choice of effective ensembles for averaging and to practically contend with the vast combinatorial complexity of mutations. In this regard, we discuss possible approaches for optimally learning the epistatic structure of biological systems.

q-bio.QM

Rediscovering the power of pairwise interactions

Two recent streams of work suggest that pairwise interactions may be sufficient to capture the complexity of biological systems ranging from protein structure to networks of neurons. In one approach, possible amino acid sequences in a family of proteins are generated by Monte Carlo annealing of a "Hamiltonian" that forces pairwise correlations among amino acid substitutions to be close to the observed correlations. In the other approach, the observed correlations among pairs of neurons are used to construct a maximum entropy model for the states of the network as a whole. We show that, in certain limits, these two approaches are mathematically equivalent, and we comment on open problems suggested by this framework

q-bio.QM