SearcharxivSearch

arXiv subjects

Francesco Zamponi

Publications and source records attributed to Francesco Zamponi.

At least 19 recordsLinked to original sources

High-dimensional theory of the glass transition revisited: hopping and local defects

The replicated liquid theory provides a microscopic mean-field description of the glass transition by combining the density functional theory of liquids with the replica method originally developed for spin glasses. In the conventional replica liquid theory, a glassy state is described by assuming that particles in different replicas undergo vibrational motion around common centers of mass, thereby forming molecules that contain one particle from every replica. Here we revisit this assumption by allowing each molecule to contain only a subset of replicas. This generalized formulation describes particle-level replica mismatches, which may be associated with non-vibrational motions such as particle hopping. We apply the theory to high-dimensional hard and harmonic spheres, where the mean-field description is expected to become exact. For hard spheres, replica mismatches destabilize the glassy metastable state and shift the dynamical transition to a significantly higher packing fraction, while leaving the leading thermodynamic glass transition unchanged. The resulting transition density agrees, at leading order in high dimensions, with the recent rigorous lower bound for random sphere packings obtained by Campos, Jenssen, Michelen, and Sahasrabudhe by using a discretized version of greedy Random Sequential Absorption, suggesting an algorithmic interpretation of the transition: grandcanonical dynamics is more efficient in high dimensional spaces than canonical one. For harmonic spheres at finite temperature, the glassy state contains a finite replica-mismatch fraction even at the thermodynamic ideal-glass transition, thereby shifting the transition point from that predicted by the conventional replica ansatz.

cond-mat.dis-nn

Modeling Protein Evolution with Generative Models: from Extant Sequence Data to Evolutionary Dynamics

Protein sequences carry a record of evolutionary history shaped by mutation, selection, drift, and epistasis. Recent generative models trained on homologous sequence families offer a new way to read this record: they define probabilistic landscapes that score sequences, generate viable variants, and capture constraints that are difficult to measure experimentally. In this review, we discuss how such landscapes can be used not only for protein design or mutation-effect prediction, but also for modeling evolutionary dynamics. We focus particularly on Direct Coupling Analysis as an interpretable and experimentally validated framework, while placing it in the broader context of generative sequence modeling. We first describe how generative sequence landscapes are inferred and assessed, then review how they can be coupled to population-genetic or substitution-model dynamics to simulate protein evolution across experimental and phylogenetic timescales. Applications include viral evolution, laboratory drift experiments, historical contingency, entrenchment, epistatic drift over time, and long-term sequence-space exploration. We conclude by discussing open challenges, including score-fitness calibration, phylogenetic structure, codon-level mutation biases, indels, and the integration of experimental data.

q-bio.PE

Towards coevolution-aware ancestral sequence reconstruction

Ancestral sequence reconstruction (ASR) is a powerful approach for studying molecular evolution and the emergence of protein function. Yet most ASR methods assume that sites evolve independently, neglecting the epistatic constraints that shape protein structure, stability, and function. This simplification affects both ancestral inference and its evaluation: maximum-a-posteriori reconstructions may over-concentrate probability into a single over-idealized sequence, whereas independent posterior sampling can generate implausible or poorly functional ancestors. Here, we introduce a coevolution-aware ASR framework that combines standard phylogenetic inference with Direct Coupling Analysis (DCA), thereby preserving site-wise ancestral uncertainty while enforcing residue-residue constraints learned from extant protein families. To benchmark the method, we develop a controlled forward-evolution framework based on a DCA evolutionary sampler, allowing reconstructed ancestors to be compared with known ground-truth sequences generated under realistic epistatic constraints. Applied to beta-lactamases and DNA-binding domains, the approach improves reconstruction when ancestral states are epistatically constrained, and yields ensembles of candidate ancestors that are both phylogenetically consistent and statistically compatible with natural protein families. This framework bridges the gap between single-sequence MAP reconstruction and unconstrained posterior sampling, providing a practical route toward ancestral reconstructions that better reflect the coupled nature of protein evolution.

q-bio.BM

A proof of an identity for the critical exponents of jamming

Within the full replica-symmetry-breaking (fullRSB) solution of dense hard spheres in infinite dimension, Charbonneau, Kurchan, Parisi, Urbani, and Zamponi (CKPUZ; J.Stat.Mech.P10009, 2014) introduced three critical exponents $a$, $b$, $c$ governing the matching region of the fullRSB profile near the jamming transition. These exponents satisfy two scaling relations. The first, $b=(1+c)/2$, was established analytically by the diffusion-drift balance in the scaling ansatz. The second, $a+b=1$, was observed numerically to arbitrary precision but could not be proven. The exponents $a,b,c$ of the scaling fullRSB ansatz are related to the physical exponents $\alpha, \theta, \kappa$ that control the gap, force, and overlap distributions by the relations $\alpha=a/b$, $\theta=(c-a)/(b-c)$, $\kappa=c+1$. Crucially, the relation $a+b=1$ yields the scaling relations $\alpha=1/(2+\theta)$ and $\kappa=2-2/(3+\theta)$ predicted on independent grounds by the mechanical-marginal-stability arguments of Wyart and collaborators. Here, we give an analytic proof of the identity $a+b=1$ from the scaling fullRSB equations. The proof was obtained through interaction with Claude (Sonnet 4.6 and Opus 4.7) and verified by us.

cond-mat.stat-mech

Expanding functional protein sequence space using high entropy generative models

Boltzmann Machines trained on evolutionary sequence data have emerged as a powerful paradigm for the data-driven design of artificial proteins. However, the relationship between model architecture, specifically parameter density, and experimental performance remains poorly understood. Here, we investigate this relationship using the Chorismate Mutase enzyme family as a model system. We compare standard fully connected Boltzmann Machines for Direct Coupling Analysis (bmDCA) with sparse models generated via progressive edge activation (eaDCA) and edge decimation (edDCA). We identify a maximum-entropy model (meDCA) along the decimation trajectory that represents an optimal balance between constraint satisfaction and the flexibility of the probability distribution. We synthesized and tested artificial sequences from all models using an in vivo complementation assay, finding that all architectures, regardless of sparsity, generate functional enzymes with high success rates, even at significant divergence from natural sequences. Despite this functional equivalence, we demonstrate that the meDCA model samples a viable sequence space that is more than fifteen orders of magnitude larger than its low-entropy counterparts. Furthermore, comparative analyses reveal that high-entropy models systematically minimize overfitting and better capture the local neutral spaces surrounding natural proteins. These findings suggest that while various models satisfying coevolutionary statistics can generate functional sequences, high-entropy Boltzmann Machines provide a superior representation of the underlying evolutionary fitness landscape.

q-bio.QM

Transition path sampling in Ising models on heterogeneous graphs

Activated transitions have rates that are often exponentially small in system size. Extracting the associated activation barriers is challenging in practice, especially in the deeply metastable regimes and in the presence of disorder. Here, we use transition path sampling to evaluate transition probabilities between ferromagnetic states in the Ising model on finite sparse random graphs, which are perhaps the simplest example of a disordered system with metastable states. To interpret the transient onset of the transition probability curve, we introduce a minimal three-state kinetic description that highlights the role of intermediate configurations. We validate the method on the heterogeneous Zachary Karate Club network, where distinct dynamical regimes emerge as temperature varies. We then apply the method to random regular graphs and Erd\H{o}s-R\'{e}nyi graphs, showing that sample-to-sample fluctuations are weak in the former but that quenched topological disorder induces sizable instance variability in the latter. For Erd\H{o}s-R\'{e}nyi graphs, we introduce an instance-dependent temperature rescaling that restores a consistent finite-size scaling of dynamical rates and enables a direct comparison with the corresponding static free-energy barrier.

cond-mat.dis-nn

Dreaming improves memorization in a Hopfield model with bounded synaptic strength

The Hopfield model provides a paradigmatic framework for associative memory. Its classical implementation, based on the Hebbian learning rule, suffers from catastrophic forgetting: when one attempts storing too many patterns, the network fails to retrieve any of them. Yet, the Hebbian rule does not take into account that synaptic strength is bounded. Introducing this biologically plausible modification, known as "clipping", eliminates catastrophic forgetting; the model is now able to retrieve the most recently seen memories, eliminating older ones. Yet, its memorization capacity is much reduced with respect to the unclipped case. Here, we investigate the effects of adding a "dreaming" phase on the capacity of a clipped Hopfield model. Following a proposal by Hopfield, Feinstein and Palmer, we assume that during the dreaming phase, the model generates random patterns that are then "unlearned". We show that while clipping still removes catastrophic forgetting, alternating learning and dreaming phases improves the memorization capacity and makes the search for optimal performance more realistic from an evolutionary perspective.

cond-mat.dis-nn

Modeling Protein Evolution via Generative Inference From Monte Carlo Chains to Population Genetics

Generative models derived from large protein sequence alignments define complex fitness landscapes, but their utility for accurately modeling non-equilibrium evolutionary dynamics remains unclear. In this work, we perform a rigorous comparative analysis of three simulation schemes, designed to mimic evolution in silico by local sampling of the probability distribution defined by a generative model. We compare standard independent Markov Chain Monte Carlo, Monte Carlo on a phylogenetic tree, and a population genetics dynamics, benchmarking their outputs against deep sequencing data from four distinct in vitro evolution experiments. We find that standard Monte Carlo fails to reproduce the correct phylogenetic structure and generates unrealistic, gradual mutational sweeps. Performing Monte Carlo on a tree inferred from data improves phylogenetic fidelity and historical accuracy. The population genetics scheme successfully captures phylogenetic correlations, mutational abundances, and selective sweeps as emergent properties, without the need to infer additional information from data. However, the latter choice come at the price of not sampling the proper generative model distribution at long times. Our findings highlight the crucial role of phylogenetic correlations and finite-population effects in shaping evolutionary trajectories on fitness landscapes. These models therefore provide powerful tools for predicting complex adaptive paths and for reliably extrapolating evolutionary dynamics beyond current experimental limitations.

q-bio.PE

Yielding in dense active matter

High-density granular active matter is a useful model for dense animal collectives and could be useful for designing reconfigurable materials that can flow or solidify on command. Recent work has demonstrated key similarities and differences between the mechanical response of dense active matter and its sheared passive counterpart, yet a constitutive law that predicts precisely how dense active matter flows or fails remains elusive. Here we study the yielding transition in dense active matter in the limit of slow driving and large persistence times, across a wide range of material preparations. Under shear, materials prepared to be very low energy or ultrastable are brittle, and well-described by elastoplastic constitutive laws. We show that under random active forcing, however, ultrastable materials are always ductile. We develop a modified elastoplastic model that captures and explains these observations, where the key parameter is the correlation length of the input active driving field. We also observe large parameter regimes where the plastic flow is surprisingly well-predicted by the input active driving field and not highly dependent on the structural disorder, suggesting new strategies for control.

cond-mat.soft

Demonstrating Real Advantage of Machine-Learning-Enhanced Monte Carlo for Combinatorial Optimization

Combinatorial optimization problems are central to both practical applications and the development of optimization methods. While classical and quantum algorithms have been refined over decades, machine learning--assisted approaches are comparatively recent and have not yet consistently outperformed simple, state-of-the-art classical methods. Here, we focus on a class of Quadratic Unconstrained Binary Optimization (QUBO) problems, specifically the challenge of finding minimum energy configurations in three-dimensional Ising spin glasses. We use a Global Annealing Monte Carlo algorithm that integrates standard local moves with global moves proposed via machine learning. We show that local moves play a crucial role in achieving optimal performance. Benchmarking against Simulated Annealing and Population Annealing, we demonstrate that Global Annealing not only surpasses the performance of Simulated Annealing but also exhibits greater robustness than Population Annealing, maintaining effectiveness across problem hardness and system size without hyperparameter tuning. These results provide clear and robust evidence that a machine learning--assisted optimization method can exceed the capabilities of classical state-of-the-art techniques in a combinatorial optimization setting.

cond-mat.dis-nn

Further testing the validity of generalized heterogeneous-elasticity theory for low-frequency excitations in structural glasses

We summarize the salient features of our theory of non-phononic vibrational excitations in glasses [W. Schirmacher et al., Nature Comm. 15, 3107 (2024)]. Next, we provide further evidence of the non-universality of the $\omega^4$ scaling of the non-phononic vibrational density of states (DoS), and the existence of an important class of non-phononic excitations in glasses, which we call defect states. These modes are induced by frozen-in stresses and can be classified as quasi-localized. Our results suggest that the commonly observed low-frequency $\omega^4$ scaling of the non-phononic vibrational density of states is highly dependent on technical aspects of the molecular dynamics simulations employed to compute the DoS.

cond-mat.dis-nn

From Scarce Functional Labels to Label-Aware Generation in Homologous Protein Families

Accurately annotating and controlling protein function from sequence data remains a major challenge in protein engineering, especially when functional labels are scarce within large homologous families. Here, we study a two-stage light-supervision strategy for fine-grained functional annotation and label-aware sequence generation. First, we compare several sequence representations, including one-hot encodings, Restricted Boltzmann Machines (RBMs), and ESM2-based protein language model embeddings, for predicting intra-family specificity labels from limited supervision. By using train/test splits that explicitly reduce phylogenetic leakage, we show that ESM2-based representations do not systematically outperform family-specific RBM embeddings or even simple one-hot baselines in this regime. Second, we use the inferred annotations to train an annotation-aware RBM capable of generating artificial homologs conditioned on prescribed labels. Across several protein families, we quantify how the number and quality of available labels determine the reliability of conditional generation. Our results show that scarce annotations can support label-aware protein design when they are accurately propagated, while also highlighting the importance of phylogeny-aware evaluation for assessing functional annotation methods within homologous families.

q-bio.QM

Performance of machine-learning-assisted Monte Carlo in sampling from simple statistical physics models

Recent years have seen a rise in the application of machine learning techniques to aid the simulation of hard-to-sample systems that cannot be studied using traditional methods. Despite the introduction of many different architectures and procedures, a wide theoretical understanding is still lacking, with the risk of suboptimal implementations. As a first step to address this gap, we provide here a complete analytic study of the widely-used Sequential Tempering procedure applied to a shallow MADE architecture for the Curie-Weiss model. The contribution of this work is twofold: firstly, we give a description of the optimal weights and of the training under Gradient Descent optimization. Secondly, we compare what happens in Sequential Tempering with and without the addition of local Metropolis Monte Carlo steps. We are thus able to give theoretical predictions on the best procedure to apply in this case. This work establishes a clear theoretical basis for the integration of machine learning techniques into Monte Carlo sampling and optimization.

cond-mat.dis-nn

Functional bottlenecks can emerge from non-epistatic underlying traits

Protein fitness landscapes frequently exhibit epistasis, where the effect of a mutation depends on the genetic context in which it occurs, i.e., the rest of the protein sequence. Epistasis increases landscape complexity, often resulting in multiple fitness peaks. In its simplest form, known as global epistasis, fitness is modeled as a non-linear function of an underlying additive trait. In contrast, more complex epistasis arises from a network of (pairwise or many-body) interactions between residues, which cannot be removed by a single non-linear transformation. Recent studies have explored how global and network epistasis contribute to the emergence of functional bottlenecks - fitness landscape topologies where two broad high-fitness basins, representing distinct phenotypes, are separated by a bottleneck that can only be crossed via one or a few mutational paths. Here, we introduce and analyze a stylized model of global epistasis with an additive underlying trait. We demonstrate that functional bottlenecks arise with high probability if the model is properly calibrated. Furthermore, our results underscore that a proper balance between neutral and non-neutral mutations is needed for the emergence of functional bottlenecks.

q-bio.PE

Rare Trajectories in a Prototypical Mean-field Disordered Model: Insights into Landscape and Instantons

For disordered systems within the random first-order transition (RFOT) universality class, such as structural glasses and certain spin glasses, the role played by activated relaxation processes is rich to the point of perplexity. Over the last decades, various efforts have attempted to formalize and systematize such processes in terms of instantons similar to the nucleation droplets of first-order phase transitions. In particular, Kirkpatrick, Thirumalai, and Wolynes proposed in the late '80s an influential nucleation theory of relaxation in structural glasses. Already within this picture, however, the resulting structures are far from the compact objects expected from the classical droplet description. In addition, an altogether different type of single-particle hopping-like instantons has recently been isolated in molecular simulations. Landscape studies of mean-field spin glass models have further revealed that simple saddle crossing does not capture relaxation in these systems. We present here a landscape-agnostic study of rare dynamical events, which delineates the richness of instantons in these systems. Our work not only captures the structure of metastable states, but also identifies the point of irreversibility, beyond which activated relaxation processes become a fait accompli. An interpretation of the associated landscape features is articulated, thus charting a path toward a complete understanding of RFOT instantons.

cond-mat.dis-nn

adabmDCA 2.0 -- a flexible but easy-to-use package for Direct Coupling Analysis

In this methods article, we provide a flexible but easy-to-use implementation of Direct Coupling Analysis (DCA) based on Boltzmann machine learning, together with a tutorial on how to use it. The package \texttt{adabmDCA 2.0} is available in different programming languages (C++, Julia, Python) usable on different architectures (single-core and multi-core CPU, GPU) using a common front-end interface. In addition to several learning protocols for dense and sparse generative DCA models, it allows to directly address common downstream tasks like residue-residue contact prediction, mutational-effect prediction, scoring of sequence libraries and generation of artificial sequences for sequence design. It is readily applicable to protein and RNA sequence data.

q-bio.QM

Fluctuations and the limit of predictability in protein evolution

Protein evolution involves mutations occurring across a wide range of time scales. In analogy with disordered systems in statistical physics, this dynamical heterogeneity suggests strong correlations between mutations happening at distinct sites and times. To quantify these correlations, we examine the role of various fluctuation sources in protein evolution, simulated using a data-driven energy landscape as a proxy for protein fitness. By applying spatio-temporal correlation functions developed in the context of disordered physical systems, we disentangle fluctuations originating from the initial condition, i.e. the ancestral sequence from which the evolutionary process originated, from those driven by stochastic mutations along independent evolutionary paths. Our analysis shows that, in diverse protein families, fluctuations from the ancestral sequence predominate at shorter time scales. This allows us to identify a time scale over which ancestral sequence information persists, enabling its reconstruction. We link this persistence to the strength of epistatic interactions: ancestral sequences with stronger epistatic signatures impact evolutionary trajectories over extended periods. At longer time scales, however, ancestral influence fades as epistatically constrained sites evolve collectively. To confirm this idea, we apply a standard ancestral sequence reconstruction algorithm and verify that the time-dependent recovery error is influenced by the properties of the ancestor itself. Overall, our results reveal that the properties of ancestral sequences - particularly their epistatic constraints - influence the initial evolutionary dynamics and the performance of standard ancestral sequence reconstruction algorithms.

q-bio.BM

Exact full-RSB SAT/UNSAT transition in infinitely wide two-layer neural networks

We analyze the problem of storing random pattern-label associations using two classes of continuous non-convex weights models, namely the perceptron with negative margin and an infinite-width two-layer neural network with non-overlapping receptive fields and generic activation function. Using a full-RSB ansatz we compute the exact value of the SAT/UNSAT transition. Furthermore, in the case of the negative perceptron we show that the overlap distribution of typical states displays an overlap gap (a disconnected support) in certain regions of the phase diagram defined by the value of the margin and the density of patterns to be stored. This implies that some recent theorems that ensure convergence of Approximate Message Passing (AMP) based algorithms to capacity are not applicable. Finally, we show that Gradient Descent is not able to reach the maximal capacity, irrespectively of the presence of an overlap gap for typical states. This finding, similarly to what occurs in binary weight models, suggests that gradient-based algorithms are biased towards highly atypical states, whose inaccessibility determines the algorithmic threshold.

cond-mat.dis-nn