SearcharxivSearch

arXiv subjects

Michele Parrinello

Publications and source records attributed to Michele Parrinello.

At least 19 recordsLinked to original sources

Let's Stalk About Membranes: Committor-Based Enhanced Sampling of Stalk Formation

Membrane fusion is essential for cellular communication and function, and understanding how two lipid bilayers merge is key to informing therapeutic strategies. Functionalized nanoparticles have recently emerged as synthetic fusogens, but the molecular mechanisms driving this process remain unclear, partly because fusion involves transitions over high free-energy barriers, difficult to capture in molecular simulations. While enhanced sampling methods can address this problem, they also rely on the definition of collective variables, which are especially hard to define for fusion, as it arises from the collective rearrangement of many molecules and cannot be easily reduced to a simple intuitive coordinate. Here, we study stalk formation, the first step of fusion, mediated by an amphiphilic gold nanoparticle, by employing an enhanced sampling strategy based on the committor function, machine-learned through a self-consistent procedure. This method requires minimal prior knowledge of the system and leverages the learned committor function as an effective collective variable, enabling uniform sampling of the entire pathway. From the resulting reactive trajectories and extensive transition region sampling, we obtain converged free-energy estimates and mechanistic insight into stalk formation.

cond-mat.soft

Ceci n'est pas un committor, yet it samples like one: efficient sampling via approximated committor functions

Atomistic simulations are widely used to investigate reactive processes but are often limited by the rare event problem due to kinetic bottlenecks. We recently introduced an enhanced sampling approach based on the committor function, machine-learned following a variational principle. This method combines a transition-state-oriented bias potential, expressed as a functional of the committor, with a metadynamics-like bias along a committor-based collective variable, enabling uniform exploration of reaction pathways. In its original formulation, the committor is represented by a neural network that takes physical descriptors as input and is trained by minimizing a functional involving gradients with respect to atomic coordinates, which can be computationally demanding in some cases. Here, we propose a simplified learning criterion formulated entirely in the descriptor space, which bypasses the need for explicit and costly coordinate gradients and provides a relaxed upper bound to the original variational principle. Although this approach does not formally target the exact committor, we show that it retains robust sampling performance while significantly reducing computational costs, thus enabling the study of processes that would be practically unfeasible using the original formulation.

physics.comp-ph

Electrochemical Interfaces at Constant Potential: Data-Efficient Transfer Learning for Machine-Learning-Based Molecular Dynamics

Simulating electrified metal/water interfaces with explicit solvent under constant potential is essential for understanding electrochemical processes, yet remains prohibitively expensive with ab initio methods. We present TRECI, a data-efficient workflow for constructing machine learning force-fields (ML-FFs) that achieve ab initio-level accuracy in electronically grand-canonical molecular dynamics. By leveraging transfer learning from general-purpose and domain-specific models, TRECI enables stable and accurate simulations across a wide potential range using a reduced number of reference configurations. This efficiency allows the use of high-level meta-GGA functionals and rigorous surface-electrification schemes. Applied to Cu(111)/water, models trained on just one thousand configurations yield accurate molecular dynamics simulations, capturing bias-dependent solvent restructuring effects not previously reported. TRECI offers a general strategy for characterising diverse materials and interfacial chemistries, significantly lowering the cost of realistic constant-potential simulations and expanding access to quantitative electrochemical modelling.

physics.comp-ph

Committors without Descriptors

The study of rare events is one of the major challenges in atomistic simulations, and several enhanced sampling methods towards its solution have been proposed. Recently, it has been suggested that the use of the committor, which provides a precise formal description of rare events, could be of use in this context. We have recently followed up on this suggestion and proposed a committor-based method that promotes frequent transitions between the metastable states of the system and allows extensive sampling of the process transition state ensemble. One of the strengths of our approach is being self-consistent and semi-automatic, exploiting a variational criterion to iteratively optimize a neural-network-based parametrization of the committor, which uses a set of physical descriptors as input. Here, we further automate this procedure by combining our previous method with the expressive power of graph neural networks, which can directly process atomic coordinates rather than descriptors. Besides applications on benchmark systems, we highlight the advantages of a graph-based approach in describing the role of solvent molecules in systems, such as ion pair dissociation or ligand binding.

physics.comp-ph

The seeds of the future are in the present: A blind exploration of metastable states

In this work, we present a novel type of molecular dynamics simulation that aims at discovering, in a blind way, new metastable states. Using only data coming from an initial unbiased simulation, and with the help of an appropriately defined loss function, we compute a bias that favors sampling yet unexplored configurational space regions, encouraging the system to leave the initial basin. In our work, we take advantage of what is normally thought to be a defect, namely the difficulty of neural networks to generalize. Contrary to most other enhanced sampling methods, which need previous knowledge of the reactive process, we are able to discover in a blind way new metastable states, overcoming otherwise insuperable kinetic bottlenecks. We illustrate the workings of the method with a number of instructive examples.

physics.chem-ph

The role of fluctuations in the nucleation process

The emergence upon cooling of an ordered solid phase from a liquid is a remarkable example of self-assembly, which has also major practical relevance. Here, we use a recently developed committor-based enhanced sampling method [Kang et al., Nat. Comput. Sci. 4, 451-460 (2024); Trizio et al., Nat. Comput. Sci. 1-10 (2025)] to explore the crystallization transition in a Lennard-Jones fluid, using Kolmogorov's variational principle. In particular, we take advantage of the properties of our sampling method to harness a large number of configurations from the transition state ensemble. From this wealth of data, we achieve precise localization of the transition state region, revealing a nucleation pathway that deviates from idealized spherical growth assumptions. Furthermore, we take advantage of the probabilistic nature of the committor to detect and analyze the fluctuations that lead to nucleation. Our study nuances classical nucleation theory by showing that the growing nucleus has a complex structure, consisting of a solid core surrounded by an interface that is more disordered than bulk liquid. We also compute from the Kolmogorov's principle a nucleation rate that is consistent with the experimental results at variance with previous computational estimates.

cond-mat.stat-mech

Weighted Active Space Protocol for Multireference Machine-Learned Potentials

Multireference methods such as multiconfiguration pair-density functional theory (MC-PDFT) offer an effective means of capturing electronic correlation in systems with significant multiconfigurational character. However, their application to train machine learning-based interatomic potentials (MLPs) for catalytic dynamics has been challenging due to the sensitivity of multireference calculations to the underlying active space, which complicates achieving consistent energies and gradients across diverse nuclear configurations. To overcome this limitation, we introduce the Weighted Active-Space Protocol (WASP), a systematic approach to assign a consistent active space for a given system across uncorrelated configurations. By integrating WASP with MLPs and enhanced sampling techniques, we propose a data-efficient active learning cycle that enables the training of an MLP on multireference data. We demonstrate the method on the TiC+-catalyzed C-H activation of methane, a reaction that poses challenges for Kohn-Sham density functional theory due to its significant multireference character. This framework enables accurate and efficient modeling of catalytic dynamics, establishing a new paradigm for simulating complex reactive processes beyond the limits of conventional electronic-structure methods.

physics.chem-ph

Fast and Fourier Features for Transfer Learning of Interatomic Potentials

Training machine learning interatomic potentials that are both computationally and data-efficient is a key challenge for enabling their routine use in atomistic simulations. To this effect, we introduce franken, a scalable and lightweight transfer learning framework that extracts atomic descriptors from pretrained graph neural networks and transfer them to new systems using random Fourier features-an efficient and scalable approximation of kernel methods. Franken enables fast and accurate adaptation of general-purpose potentials to new systems or levels of quantum mechanical theory without requiring hyperparameter tuning or architectural modifications. On a benchmark dataset of 27 transition metals, franken outperforms optimized kernel-based methods in both training time and accuracy, reducing model training from tens of hours to minutes on a single GPU. We further demonstrate the framework's strong data-efficiency by training stable and accurate potentials for bulk water and the Pt(111)/water interface using just tens of training structures. Our open-source implementation (https://franken.readthedocs.io) offers a fast and practical solution for training potentials and deploying them for molecular dynamics simulations across diverse systems.

physics.comp-ph

The need to implement FAIR principles in biomolecular simulations

This letter illustrates the opinion of the molecular dynamics (MD) community on the need to adopt a new FAIR paradigm for the use of molecular simulations. It highlights the necessity of a collaborative effort to create, establish, and sustain a database that allows findability, accessibility, interoperability, and reusability of molecular dynamics simulation data. Such a development would democratize the field and significantly improve the impact of MD simulations on life science research. This will transform our working paradigm, pushing the field to a new frontier. We invite you to support our initiative at the MDDB community (https://mddbr.eu/community/) Now published as: Amaro, R.E., et al. The need to implement FAIR principles in biomolecular simulations. Nat Methods (2025) https://doi.org/10.1038/s41592-025-02635-0

q-bio.BM

Everything everywhere all at once: a probability-based enhanced sampling approach to rare events

The problem of studying rare events is central to many areas of computer simulations. In a recent paper [Kang, P., et al., Nat. Comput. Sci. 4, 451-460, 2024], we have shown that a powerful way of solving this problem passes through the computation of the committor function, and we have demonstrated how the committor can be iteratively computed in a variational way and the transition state ensemble efficiently sampled. Here, we greatly ameliorate this procedure by combining it with a metadynamics-like enhanced sampling approach in which a logarithmic function of the committor is used as a collective variable. This integrated procedure leads to an accurate and balanced sampling of the free energy surface in which transition states and metastable basins are studied with the same thoroughness. We also show that our approach can be used in cases in which competing reactive paths are possible and intermediate metastable are encountered. In addition, we demonstrate how physical insights can be obtained from the optimized committor model and the sampled data, thus providing a full characterization of the rare event under study. We ascribe the success of this approach to the use of a probability-based description of rare events.

physics.comp-ph

Descriptors-free Collective Variables From Geometric Graph Neural Networks

Enhanced sampling simulations make the computational study of rare events feasible. A large family of such methods crucially depends on the definition of some collective variables (CVs) that could provide a low-dimensional representation of the relevant physics of the process. Recently, many methods have been proposed to semi-automatize the CV design by using machine learning tools to learn the variables directly from the simulation data. However, most methods are based on feed-forward neural networks and require as input some user-defined physical descriptors. Here, we propose to bypass this step using a graph neural network to directly use the atomic coordinates as input for the CV model. This way, we achieve a fully automatic approach to CV determination that provides variables invariant under the relevant symmetries, especially the permutational one. Furthermore, we provide different analysis tools to favor the physical interpretation of the final CV. We prove the robustness of our approach using different methods from the literature for the optimization of the CV, and we prove its efficacy on several systems, including a small peptide, an ion dissociation in explicit solvent, and a simple chemical reaction.

physics.comp-ph

From Biased to Unbiased Dynamics: An Infinitesimal Generator Approach

We investigate learning the eigenfunctions of evolution operators for time-reversal invariant stochastic processes, a prime example being the Langevin equation used in molecular dynamics. Many physical or chemical processes described by this equation involve transitions between metastable states separated by high potential barriers that can hardly be crossed during a simulation. To overcome this bottleneck, data are collected via biased simulations that explore the state space more rapidly. We propose a framework for learning from biased simulations rooted in the infinitesimal generator of the process and the associated resolvent operator. We contrast our approach to more common ones based on the transfer operator, showing that it can provably learn the spectral properties of the unbiased system from biased data. In experiments, we highlight the advantages of our method over transfer operator approaches and recent developments based on generator learning, demonstrating its effectiveness in estimating eigenfunctions and eigenvalues. Importantly, we show that even with datasets containing only a few relevant transitions due to sub-optimal biasing, our approach recovers relevant information about the transition mechanism.

cs.LG

PLUMED Tutorials: a collaborative, community-driven learning ecosystem

In computational physics, chemistry, and biology, the implementation of new techniques in a shared and open source software lowers barriers to entry and promotes rapid scientific progress. However, effectively training new software users presents several challenges. Common methods like direct knowledge transfer and in-person workshops are limited in reach and comprehensiveness. Furthermore, while the COVID-19 pandemic highlighted the benefits of online training, traditional online tutorials can quickly become outdated and may not cover all the software's functionalities. To address these issues, here we introduce ``PLUMED Tutorials'', a collaborative model for developing, sharing, and updating online tutorials. This initiative utilizes repository management and continuous integration to ensure compatibility with software updates. Moreover, the tutorials are interconnected to form a structured learning path and are enriched with automatic annotations to provide broader context. This paper illustrates the development, features, and advantages of PLUMED Tutorials, aiming to foster an open community for creating and sharing educational resources.

physics.ed-ph

DLGNet: Hyperedge Classification through Directed Line Graphs for Chemical Reactions

Graphs and hypergraphs provide powerful abstractions for modeling interactions among a set of entities of interest and have been attracting a growing interest in the literature thanks to many successful applications in several fields. In particular, they are rapidly expanding in domains such as chemistry and biology, especially in the areas of drug discovery and molecule generation. One of the areas witnessing the fasted growth is the chemical reactions field, where chemical reactions can be naturally encoded as directed hyperedges of a hypergraph. In this paper, we address the chemical reaction classification problem by introducing the notation of a Directed Line Graph (DGL) associated with a given directed hypergraph. On top of it, we build the Directed Line Graph Network (DLGNet), the first spectral-based Graph Neural Network (GNN) expressly designed to operate on a hypergraph via its DLG transformation. The foundation of DLGNet is a novel Hermitian matrix, the Directed Line Graph Laplacian, which compactly encodes the directionality of the interactions taking place within the directed hyperedges of the hypergraph thanks to the DLG representation. The Directed Line Graph Laplacian enjoys many desirable properties, including admitting an eigenvalue decomposition and being positive semidefinite, which make it well-suited for its adoption within a spectral-based GNN. Through extensive experiments on chemical reaction datasets, we show that DGLNet significantly outperforms the existing approaches, achieving on a collection of real-world datasets an average relative-percentage-difference improvement of 33.01%, with a maximum improvement of 37.71%.

cs.LG

Effective Data-Driven Collective Variables for Free Energy Calculations from Metadynamics of Paths

A variety of enhanced sampling methods predict multidimensional free energy landscapes associated with biological and other molecular processes as a function of a few selected collective variables (CVs). The accuracy of these methods is crucially dependent on the ability of the chosen CVs to capture the relevant slow degrees of freedom of the system. For complex processes, finding such CVs is the real challenge. Machine learning (ML) CVs offer, in principle, a solution to handle this problem. However, these methods rely on the availability of high-quality datasets -- ideally incorporating information about physical pathways and transition states -- which are difficult to access, therefore greatly limiting their domain of application. Here, we demonstrate how these datasets can be generated by means of enhanced sampling simulations in trajectory space via the metadynamics of paths [arXiv:2002.09281] algorithm. The approach is expected to provide a general and efficient way to generate efficient ML-based CVs for the fast prediction of free energy landscapes in enhanced sampling simulations. We demonstrate our approach with two numerical examples, a two-dimensional model potential and the isomerization of alanine dipeptide, using deep targeted discriminant analysis as our ML-based CV of choice.

physics.comp-ph

Computing the Committor with the Committor: an Anatomy of the Transition State Ensemble

Determining the kinetic bottlenecks that make transitions between metastable states difficult is key to understanding important physical problems like crystallization, chemical reactions, or protein folding. In all these phenomena, the system spends a considerable amount of time in one metastable state before making a rare but important transition to a new state. The rarity of these events makes their direct simulation challenging, if not impossible. We propose a method to explore the distribution of configurations that the system passes as it translocates from one metastable basin to another. We shall refer to this set of configurations as the transition state ensemble. We base our method on the committor function and the variational principle to which it obeys. We find the minimum of the variational principle via a self-consistent procedure that does not require any input besides the knowledge of the initial and final state. Right from the start, our procedure focuses on sampling the transition state ensemble and allows harnessing a large number of such configurations. With the help of the variational principle, we perform a detailed analysis of the transition state ensemble, ranking quantitatively the degrees of freedom mostly involved in the transition and opening the way for a systematic approach for the interpretation of simulation results and the construction of collective variables.

physics.comp-ph

Unraveling the Crystallization Kinetics of the Ge$_2$Sb$_2$Te$_5$ Phase Change Compound with a Machine-Learned Interatomic Potential

The phase change compound Ge$_2$Sb$_2$Te$_5$ (GST225) is exploited in advanced non-volatile electronic memories and in neuromorphic devices which both rely on a fast and reversible transition between the crystalline and amorphous phases induced by Joule heating. The crystallization kinetics of GST225 is a key functional feature for the operation of these devices. We report here on the development of a machine-learned interatomic potential for GST225 that allowed us to perform large scale molecular dynamics simulations (over 10000 atoms for over 100 ns) to uncover the details of the crystallization kinetics in a wide range of temperatures of interest for the programming of the devices. The potential is obtained by fitting with a deep neural network (NN) scheme a large quantum-mechanical database generated within Density Functional Theory. The availability of a highly efficient and yet highly accurate NN potential opens the possibility to simulate phase change materials at the length and time scales of the real devices.

cond-mat.mtrl-sci

Transfer learning for atomistic simulations using GNNs and kernel mean embeddings

Interatomic potentials learned using machine learning methods have been successfully applied to atomistic simulations. However, accurate models require large training datasets, while generating reference calculations is computationally demanding. To bypass this difficulty, we propose a transfer learning algorithm that leverages the ability of graph neural networks (GNNs) to represent chemical environments together with kernel mean embeddings. We extract a feature map from GNNs pre-trained on the OC20 dataset and use it to learn the potential energy surface from system-specific datasets of catalytic processes. Our method is further enhanced by incorporating into the kernel the chemical species information, resulting in improved performance and interpretability. We test our approach on a series of realistic datasets of increasing complexity, showing excellent generalization and transferability performance, and improving on methods that rely on GNNs or ridge regression alone, as well as similar fine-tuning approaches.

cs.LG