SearcharxivSearch

arXiv subjects

Thierry Mora

Publications and source records attributed to Thierry Mora.

At least 19 recordsLinked to original sources

Surveying the adaptive landscapes of 10,000 antibodies

Affinity maturation is the Darwinian process by which antibodies improve antigen binding through somatic hypermutation and selection. The adaptive landscape, which defines the set of antibody-specific mutations that improve functional characteristics like antigen binding, has been explored in only a handful of antibodies. Identifying the sites of adaptive mutations in a given antibody sequence, and how these sites vary across the antibody repertoire, can inform the design of therapeutic antibodies. We develop a parameter-free population genetic framework that leverages the statistics of convergent affinity maturation in B cell lineages sharing similar naive sequences, called public clonotypes, to identify beneficial mutations. Applying this framework to more than 10,000 public clonotypes represented by multiple lineages across 20 healthy individuals, we identify widespread signatures of clonotype-dependent selection of individual mutations. We estimate the prevalence and typical fitness effects of mutations across the V gene at the single-site level, uncovering a general tradeoff between prevalence and fitness effect. These inferred landscapes broadly reproduce the statistics of convergent mutation in antibodies specific to SARS-CoV-2 and influenza. Finally, we use our framework to benchmark predictions from existing antibody language models, and show that while these models are dominated by non-selective signatures, a simple renormalization procedure can expose signatures of clonotype-dependent positive selection consistent with our predictions.

q-bio.PE

Neutralization titers reveal the structure of polyclonal antibody responses

The composition of a polyclonal antibody response is hard to measure experimentally but contains vital information about the robustness of immunity. Here, we argue that the statistics of neutralization titers alone can be used to make quantitative predictions about the composition of the response, circumventing challenges arising through sequencing and monoclonal antibody expression. We show that the response against influenza within a cohort can be either driven by a collective phenomenon where many antibodies contribute to neutralization, or dominated by just a few strong binders, leading to a broad distribution of titers across individuals described by a Gumbel distribution from extreme value theory. Comparing titers across cohorts, we find that Gumbel statistics {accurately describe} individuals prior to an immune challenge. We propose an equilibrium binding model that quantitatively captures titer data and illustrates the structure of the polyclonal response. Our approach extends generically to immune responses to other pathogens.

q-bio.PE

Modeling Protein Evolution via Generative Inference From Monte Carlo Chains to Population Genetics

Generative models derived from large protein sequence alignments define complex fitness landscapes, but their utility for accurately modeling non-equilibrium evolutionary dynamics remains unclear. In this work, we perform a rigorous comparative analysis of three simulation schemes, designed to mimic evolution in silico by local sampling of the probability distribution defined by a generative model. We compare standard independent Markov Chain Monte Carlo, Monte Carlo on a phylogenetic tree, and a population genetics dynamics, benchmarking their outputs against deep sequencing data from four distinct in vitro evolution experiments. We find that standard Monte Carlo fails to reproduce the correct phylogenetic structure and generates unrealistic, gradual mutational sweeps. Performing Monte Carlo on a tree inferred from data improves phylogenetic fidelity and historical accuracy. The population genetics scheme successfully captures phylogenetic correlations, mutational abundances, and selective sweeps as emergent properties, without the need to infer additional information from data. However, the latter choice come at the price of not sampling the proper generative model distribution at long times. Our findings highlight the crucial role of phylogenetic correlations and finite-population effects in shaping evolutionary trajectories on fitness landscapes. These models therefore provide powerful tools for predicting complex adaptive paths and for reliably extrapolating evolutionary dynamics beyond current experimental limitations.

q-bio.PE

Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics

The specific region of an antibody responsible for binding to an antigen, known as the paratope, is essential for immune recognition. Accurate identification of this small yet critical region can accelerate the development of therapeutic antibodies. Determining paratope locations typically relies on modeling the antibody structure, which is computationally intensive and difficult to scale across large antibody repertoires. We introduce Paraplume, a sequence-based paratope prediction method that leverages embeddings from protein language models (PLMs), without the need for structural input and achieves superior performance across multiple benchmarks compared to current methods. In addition, reweighting PLM embeddings using Paraplume predictions yields more informative sequence representations, improving downstream tasks such as affinity prediction, binder classification, and epitope binning. Applied to large antibody repertoires, Paraplume reveals that antigen-specific somatic hypermutations are associated with larger paratopes, suggesting a potential mechanism for affinity enhancement. Our findings position PLM-based paratope prediction as a powerful, scalable alternative to structure-dependent approaches, opening new avenues for understanding antibody evolution.

q-bio.BM

Dynamics of memory B cells and plasmablasts in healthy individuals

Our adaptive immune system relies on the persistence over long times of a diverse set of antigen-experienced B cells to encode our memories of past infections and to protect us against future ones. While longitudinal repertoire sequencing promises to track the long-term dynamics of many B cell clones simultaneously, sampling and experimental noise make it hard to draw reliable quantitative conclusions. Leveraging statistical inference, we infer the dynamics of memory B cell clonal dynamics and conversion to plasmablasts, which includes clone creation, degradation, abundance fluctuations, and differentiation. We find that memory B cell clones degrade slowly, with a half-life of 10 years. Based on the inferred parameters, we predict that it takes about 50 years to renew 50\% of the repertoire, with most observed clones surviving for a lifetime. We infer that, on average, 1 out of 100 memory B cells differentiates into a plasmablast each year, more than expected from purely antigen-stimulated differentiation, and that plasmablast clones degrade with a half-life of about one year in the absence of memory imports. Our method is general and could be applied to other longitudinal repertoire sequencing B cell subsets.

q-bio.PE

Theoretical limits for sensing through phase separation

Biomolecular condensates form on timescales of seconds in cells upon environmental or compositional changes. Condensate formation is thus argued to act as a mechanism for sensing such changes and quickly initiating downstream processes, such as forming stress granules in response to heat stress and amplifying cGAS enzymatic activity upon detection of cytosolic DNA. Here, we study a dynamical model of droplet nucleation and growth to demonstrate how phase separation allows cells to discriminate small concentration differences on finite, biologically relevant timescales. We propose optimal sensing protocols, which use the sharp onset of phase separation. We show how, given experimentally measured rates, cells can achieve rapid and robust sensing of concentration differences of 1% on a timescale of minutes, offering an alternative to classical biochemical mechanisms.

q-bio.SC

Fast decisions with biophysically constrained gene promoter architectures

Cells integrate signals and make decisions about their future state in short amounts of time. A lot of theoretical effort has gone into asking how to best design gene regulatory circuits that fulfill a given function, yet little is known about the constraints that performing that function in a small amount of time imposes on circuit architectures. Using an optimization framework, we explore the properties of a class of promoter architectures that distinguish small differences in transcription factor concentrations under time constraints. We show that the full temporal trajectory of gene activity allows for faster decisions than its integrated activity represented by the total number of transcribed mRNA. The topology of promoter architectures that allow for rapidly distinguishing low transcription factor concentrations result in a low, shallow, and non cooperative response, while at high concentrations, the response is high and cooperative. In the presence of non-cognate ligands, networks with fast and accurate decision times need not be optimally selective, especially if discrimination is difficult. While optimal networks are generically out of equilibrium, the energy associated with that irreversibility is only modest, and negligible at small concentrations. Instead, our results highlight the crucial role of rate-limiting steps imposed by biophysical constraints.

q-bio.MN

Integrating computational detection and experimental validation for rapid GFRAL-specific antibody discovery

The identification and validation of therapeutic antibodies is critical for developing effective treatments for many diseases. We present a computational approach for identifying antibodies targeting GFRAL-specific receptors, receptors implicated in appetite regulation. Using humanized Trianni mice, we conducted a longitudinal study with repeated blood sampling and splenic analysis. We applied the STAR computational method for antibody discovery on bulk antibody repertoire data sampled at key time points. By mapping the output from STAR to single-cell data taken at the last time point, we successfully identified a pool of antibodies, of which 50% demonstrated binding capabilities. We observed convergent selection, where responding sequences with identical amino acid complementarity determining regions 3 (CDR3) were found in different mice. We provide a catalog of 67 experimentally validated antibodies against GFRAL. The potential of these antibodies as antagonists or agonists against GFRAL suggests therapeutic solutions for conditions like cancer cachexia, anorexia, obesity, and diabetes. This study underscores the utility of integrating computational methods and experimental validation for antibody discovery in therapeutic contexts by reducing time and increasing efficiency.

q-bio.TO

Learning via mechanosensitivity and activity in cytoskeletal networks

In this work we show how a network inspired by a coarse-grained description of actomyosin cytoskeleton can learn - in a contrastive learning framework - from environmental perturbations if it is endowed with mechanosensitive proteins and motors. Our work is a proof of principle for how force-sensitive proteins and molecular motors can form the basis of a general strategy to learn in biological systems. Our work identifies a minimal biologically plausible learning mechanism and also explores its implications for commonly occuring phenomenolgy such as adaptation and homeostatis.

cond-mat.soft

Inferring resource competition in microbial communities from time series

The competition for resources is a defining feature of microbial communities. In many contexts, from soils to host-associated communities, highly diverse microbes are organized into metabolic groups or guilds with similar resource preferences. The resource preferences of individual taxa that give rise to these guilds are critical for understanding fluxes of resources through the community and the structure of diversity in the system. However, inferring the metabolic capabilities of individual taxa, and their competition with other taxa, within a community is challenging and unresolved. Here we address this gap in knowledge by leveraging dynamic measurements of abundances in communities. We show that simple correlations are often misleading in predicting resource competition. We show that spectral methods such as the cross-power spectral density (CPSD) and coherence that account for time-delayed effects are superior metrics for inferring the structure of resource competition in communities. We first demonstrate this fact on synthetic data generated from consumer-resource models with time-dependent resource availability, where taxa are organized into groups or guilds with similar resource preferences. By applying spectral methods to oceanic plankton time-series data, we demonstrate that these methods detect interaction structures among species with similar genomic sequences. Our results indicate that analyzing temporal data across multiple timescales can reveal the underlying structure of resource competition within communities.

physics.soc-ph

Energy-based generative models for monoclonal antibodies

Since the approval of the first antibody drug in 1986, a total of 162 antibodies have been approved for a wide range of therapeutic areas, including cancer, autoimmune, infectious, or cardiovascular diseases. Despite advances in biotechnology that accelerated the development of antibody drugs, the drug discovery process for this modality remains lengthy and costly, requiring multiple rounds of optimizations before a drug candidate can progress to preclinical and clinical trials. This multi-optimization problem involves increasing the affinity of the antibody to the target antigen while refining additional biophysical properties that are essential to drug development such as solubility, thermostability or aggregation propensity. Additionally, antibodies that resemble natural human antibodies are particularly desirable, as they are likely to offer improved profiles in terms of safety, efficacy, and reduced immunogenicity, further supporting their therapeutic potential. In this article, we explore the use of energy-based generative models to optimize a candidate monoclonal antibody. We identify tradeoffs when optimizing for multiple properties, concentrating on solubility, humanness and affinity and use the generative model we develop to generate candidate antibodies that lie on an optimal Pareto front that satisfies these constraints.

q-bio.BM

Single cells can resolve graded stimuli

Cells use signalling pathways as windows into the environment to gather information, transduce it into their interior, and use it to drive behaviours. MAPK (ERK) is a highly conserved signalling pathway in eukaryotes, directing multiple fundamental cellular behaviours such as proliferation, migration, and differentiation, making it of few central hubs in the signalling circuitry of cells. Despite this versatility of behaviors, population-level measurements have reported low information content (\textless 1 bit) relayed through the ERK pathway, rendering the population barely able to distinguish the presence or absence of stimuli. Here, we contrast the information transmitted by a single cell and a population of cells. Using a combination of optogenetic experiments, data analysis based on information theory framework, and numerical simulations we quantify the amount of information transduced from the receptor to ERK, from responses to singular, brief and sparse input pulses. We show that single cells are indeed able to resolve between graded stimuli, yielding over 2 bit of information, however showing a large population heterogeneity.

q-bio.CB

How host mobility patterns shape antigenic escape during viral-immune co-evolution

Viruses like influenza have long coevolved with host immune systems, gradually shaping the evolutionary trajectory of these pathogens. Host immune systems develop immunity against circulating strains, which in turn avoid extinction by exploiting antigenic escape mutations that render new strains immune from existing antibodies in the host population. Infected hosts are also mobile, which can spread the virus to regions without developed host immunity, offering additional reservoirs for viral growth. While the effects of migration on long term stability have been investigated, we know little about how antigenic escape coupled with migration changes the survival and spread of emerging viruses. By considering the two processes on equal footing, we show that on short timescales an intermediate host mobility rate increases the survival probability of the virus through antigenic escape. We show that more strongly connected migratory networks decrease the survival probability of the virus. Using data from high traffic airports we argue that current human migration rates are beneficial for viral survival.

q-bio.PE

Strong, but not weak, noise correlations are beneficial for population coding

Neural correlations play a critical role in sensory information coding. They are of two kinds: signal correlations, when neurons have overlapping sensitivities, and noise correlations from network effects and shared noise. In experiments from early sensory systems and cortex, many pairs of neurons typically show both types of correlations to be positive and large, especially between nearby neurons with similar stimulus sensitivity. However, theoretical arguments have suggested that stimulus and noise correlations should have opposite signs to improve coding, at odds with experimental observations. We analyze retinal recording in response to a large variety of stimuli, and show that, contrary to common belief, large noise correlations are beneficial for coding, even if aligned with signal correlations. To understand this result, we develop a theory of visual information coding by correlated neurons, which resolves that paradox. We show that noise correlations are always beneficial if they are strong enough, unless neurons are perfectly correlated by the stimulus. Finally, using neuronal recordings and modeling, we show that for high dimensional stimuli noise correlation benefits the encoding of fine-grained details of visual stimuli, at the expense of large-scale features, which are already well encoded.

q-bio.NC

Combining mutation and recombination statistics to infer clonal families in antibody repertoires

B-cell repertoires are characterized by a diverse set of receptors of distinct specificities generated through two processes of somatic diversification: V(D)J recombination and somatic hypermutations. B cell clonal families stem from the same V(D)J recombination event, but differ in their hypermutations. Clonal families identification is key to understanding B-cell repertoire function, evolution and dynamics. We present HILARy (High-precision Inference of Lineages in Antibody Repertoires), an efficient, fast and precise method to identify clonal families from single- or paired-chain repertoire sequencing datasets. HILARy combines probabilistic models that capture the receptor generation and selection statistics with adapted clustering methods to achieve consistently high inference accuracy. It automatically leverages the phylogenetic signal of shared mutations in difficult repertoire subsets. Exploiting the high sensitivity of the method, we find the statistics of evolutionary properties such as the site frequency spectrum and dN/dS ratio do not depend on the junction length. We also identify a broad range of selection pressures spanning two orders of magnitude.

q-bio.PE

Computational detection of antigen specific B cell receptors following immunization

B cell receptors (BCRs) play a crucial role in recognizing and fighting foreign antigens. High-throughput sequencing enables in-depth sampling of the BCRs repertoire after immunization. However, only a minor fraction of BCRs actively participate in any given infection. To what extent can we accurately identify antigen-specific sequences directly from BCRs repertoires? We present a computational method grounded on sequence similarity, aimed at identifying statistically significant responsive BCRs. This method leverages well-known characteristics of affinity maturation and expected diversity. We validate its effectiveness using longitudinally sampled human immune repertoire data following influenza vaccination and Sars-CoV-2 infections. We show that different lineages converge to the same responding CDR3, demonstrating convergent selection within an individual. The outcomes of this method hold promise for application in vaccine development, personalized medicine, and antibody-derived therapeutics.

q-bio.PE

Variability in the local and global composition of human T-cell receptor repertoires during thymic development across cell types and individuals

The adaptive immune response relies on T cells that combine phenotypic specialization with diversity of T cell receptors (TCRs) to recognize a wide range of pathogens. TCRs are acquired and selected during T cell maturation in the thymus. Characterizing TCR repertoires across individuals and T cell maturation stages is important for better understanding adaptive immune responses and for developing new diagnostics and therapies. Analyzing a dataset of human TCR repertoires from thymocyte subsets, we find that the variability between individuals generated during the TCR V(D)J recombination is maintained through all stages of T cell maturation and differentiation. The inter-individual variability of repertoires of the same cell type is of comparable magnitude to the variability across cell types within the same individual. To zoom in on smaller scales than whole repertoires, we defined a distance measuring the relative overlap of locally similar sequences in repertoires. We find that the whole repertoire models correctly predict local similarity networks, suggesting a lack of forbidden T cell receptor sequences. The local measure correlates well with distances calculated using whole repertoire traits and carries information about cell types.

q-bio.QM

Generalized Glauber dynamics for inference in biology

Large interacting systems in biology often exhibit emergent dynamics, such as coexistence of multiple time scales, manifested by fat tails in the distribution of waiting times. While existing tools in statistical inference, such as maximum entropy models, reproduce the empirical steady state distributions, it remains challenging to learn dynamical models. We present a novel inference method, called generalized Glauber dynamics. Constructed through a non-Markovian fluctuation dissipation theorem, generalized Glauber dynamics tunes the dynamics of an interacting system, while keeping the steady state distribution fixed. We motivate the need for the method on real data from Eco-HAB, an automated habitat for testing behavior in groups of mice under semi-naturalistic conditions, and present it on simple Ising spin systems. We show its applicability for experimental data, by inferring dynamical models of social interactions in a group of mice that reproduce both its collective behavior and the long tails observed in individual dynamics.

physics.bio-ph