SearcharxivSearch

arXiv subjects

Paolo Carloni

Publications and source records attributed to Paolo Carloni.

18 recordsLinked to original sources

Extreme scaling of the metadynamics of paths algorithm on the pre-exascale JUWELS Booster supercomputer

Molecular dynamics (MD)-based path sampling algorithms are a very important class of methods used to study the energetics and kinetics of rare (bio)molecular events. They sample the highly informative but highly unlikely reactive trajectories connecting different metastable states of complex (bio)molecular systems. The metadynamics of paths (MoP) method proposed by Mandelli, Hirshberg, and Parrinello [Pys. Rev. Lett. 125 2, 026001 (2020)] is based on the Onsager-Machlup path integral formalism. This provides an analytical expression for the probability of sampling stochastic trajectories of given duration. In practice, the method samples reactive paths via metadynamics simulations performed directly in the phase space of all possible trajectories. Its parallel implementation is in principle infinitely scalable, allowing arbitrarily long trajectories to be simulated. Paving the way for future applications to study the thermodynamics and kinetics of protein-ligand (un)binding, a problem of great pharmaceutical interest, we present here the efficient implementation of MoP in the HPC-oriented biomolecular simulation software GROMACS. Our benchmarks on a membrane protein (150,000 atoms) show an unprecedented weak scaling parallel efficiency of over 70% up to 3200 GPUs on the pre-exascale JUWELS Booster machine at the Jülich Supercomputing Center.

physics.comp-ph

The need to implement FAIR principles in biomolecular simulations

This letter illustrates the opinion of the molecular dynamics (MD) community on the need to adopt a new FAIR paradigm for the use of molecular simulations. It highlights the necessity of a collaborative effort to create, establish, and sustain a database that allows findability, accessibility, interoperability, and reusability of molecular dynamics simulation data. Such a development would democratize the field and significantly improve the impact of MD simulations on life science research. This will transform our working paradigm, pushing the field to a new frontier. We invite you to support our initiative at the MDDB community (https://mddbr.eu/community/) Now published as: Amaro, R.E., et al. The need to implement FAIR principles in biomolecular simulations. Nat Methods (2025) https://doi.org/10.1038/s41592-025-02635-0

q-bio.BM

MiMiC: A High-Performance Framework for Multiscale Molecular Dynamics Simulations

MiMiC is a framework for performing multiscale simulations in which loosely coupled external programs describe individual subsystems at different resolutions and levels of theory. To make it highly efficient and flexible, we adopt an interoperable approach based on a multiple-program multiple-data (MPMD) paradigm, serving as an intermediary responsible for fast data exchange and interactions between the subsystems. The main goal of MiMiC is to avoid interfering with the underlying parallelization of the external programs, including the operability on hybrid architectures (e.g., CPU/GPU), and keep their setup and execution as close as possible to the original. At the moment, MiMiC offers an efficient implementation of electrostatic embedding QM/MM that has demonstrated unprecedented parallel scaling in simulations of large biomolecules using CPMD and GROMACS as QM and MM engines, respectively. However, as it is designed for high flexibility with general multiscale models in mind, it can be straightforwardly extended beyond QM/MM. In this article, we illustrate the software design and the features of the framework, which make it a compelling choice for multiscale simulations in the upcoming era of exascale high-performance computing.

physics.chem-ph

Effective Data-Driven Collective Variables for Free Energy Calculations from Metadynamics of Paths

A variety of enhanced sampling methods predict multidimensional free energy landscapes associated with biological and other molecular processes as a function of a few selected collective variables (CVs). The accuracy of these methods is crucially dependent on the ability of the chosen CVs to capture the relevant slow degrees of freedom of the system. For complex processes, finding such CVs is the real challenge. Machine learning (ML) CVs offer, in principle, a solution to handle this problem. However, these methods rely on the availability of high-quality datasets -- ideally incorporating information about physical pathways and transition states -- which are difficult to access, therefore greatly limiting their domain of application. Here, we demonstrate how these datasets can be generated by means of enhanced sampling simulations in trajectory space via the metadynamics of paths [arXiv:2002.09281] algorithm. The approach is expected to provide a general and efficient way to generate efficient ML-based CVs for the fast prediction of free energy landscapes in enhanced sampling simulations. We demonstrate our approach with two numerical examples, a two-dimensional model potential and the isomerization of alanine dipeptide, using deep targeted discriminant analysis as our ML-based CV of choice.

physics.comp-ph

Multimap targeted free energy estimation

We present a new method to compute free energies at a quantum mechanical (QM) level of theory from molecular simulations using cheap reference potential energy functions, such as force fields. To overcome the poor overlap between the reference and target distributions, we generalize targeted free energy perturbation (TFEP) to employ multiple configuration maps. While TFEP maps have been obtained before from an expensive training of a normalizing flow neural network (NN), our multimap estimator allows us to use the same set of QM calculations to both optimize the maps and estimate the free energy, thus removing almost completely the overhead due to training. A multimap extension of the multistate Bennett acceptance ratio estimator is also derived for cases where samples from two or more states are available. Furthermore, we propose a one-epoch learning policy that can be used to efficiently avoid overfitting when computing the loss function is expensive compared to generating data. Finally, we show how our multimap approach can be combined with enhanced sampling strategies to overcome the pervasive problem of poor convergence due to slow degrees of freedom. We test our method on the HiPen dataset of drug-like molecules and fragments, and we show that it can accelerate the calculation of the free energy difference of switching from a force field to a DFTB3 potential by about 3 orders of magnitude compared to standard FEP and by a factor of about 8 compared to previously published nonequilibrium calculations.

physics.comp-ph

Scalability of 3D-DFT by block tensor-matrix multiplication on the JUWELS Cluster

The 3D Discrete Fourier Transform (DFT) is a technique used to solve problems in disparate fields. Nowadays, the commonly adopted implementation of the 3D-DFT is derived from the Fast Fourier Transform (FFT) algorithm. However, evidence indicates that the distributed memory 3D-FFT algorithm does not scale well due to its use of all-to-all communication. Here, building on the work of Sedukhin \textit{et al}. [Proceedings of the 30th International Conference on Computers and Their Applications, CATA 2015 pp. 193-200 (01 2015)], we revisit the possibility of improving the scaling of the 3D-DFT by using an alternative approach that uses point-to-point communication, albeit at a higher arithmetic complexity. The new algorithm exploits tensor-matrix multiplications on a volumetrically decomposed domain via three specially adapted variants of Cannon's algorithm. It has here been implemented as a C++ library called S3DFT and tested on the JUWELS Cluster at the Jülich Supercomputing Center. Our implementation of the shared memory tensor-matrix multiplication attained 88\% of the theoretical single node peak performance. One variant of the distributed memory tensor-matrix multiplication shows excellent scaling, while the other two show poorer performance, which can be attributed to their intrinsic communication patterns. A comparison of S3DFT with the Intel MKL and FFTW3 libraries indicates that currently iMKL performs best overall, followed in order by FFTW3 and S3DFT. This picture might change with further improvements of the algorithm and/or when running on clusters that use network connections with higher latency, e.g. on cloud platforms.

physics.comp-ph

Targeted free energy perturbation revisited: Accurate free energies from mapped reference potentials

We present an approach that extends the theory of targeted free energy perturbation (TFEP) to calculate free energy differences and free energy surfaces at an accurate quantum mechanical level of theory from a cheaper reference potential. The convergence is accelerated by a mapping function that increases the overlap between the target and the reference distributions. Building on recent work, we show that this map can be learned with a normalizing flow neural network, without requiring simulations with the expensive target potential but only a small number of single-point calculations, and, crucially, avoiding the systematic error that was found previously. We validate the method by numerically evaluating the free energy difference in a system with a double-well potential and by describing the free energy landscape of a simple chemical reaction in the gas phase.

physics.comp-ph

Exhaustive Search of Ligand Binding Pathways via Volume-based Metadynamics

Determining the complete set of ligands' binding/unbinding pathways is important for drug discovery and to rationally interpret mutation data. Here we have developed a metadynamics-based technique that addressed this issue and allows estimating affinities in the presence of multiple escape pathways. Our approach is shown on a Lysozyme T4 variant in complex with the benzene molecule. The calculated binding free energy is in agreement with experimental data. Remarkably, not only we were able to find all the previously identified ligand binding pathways, but also we uncovered 3 new ones. This results were obtained at a small computational cost, making this approach valuable for practical applications, such as screening of small compounds libraries.

physics.chem-ph

Open Boundary Simulations of Proteins and Their Hydration Shells by Hamiltonian Adaptive Resolution Scheme

The recently proposed Hamiltonian Adaptive Resolution Scheme (H-AdResS) allows to perform molecular simulations in an open boundary framework. It allows to change on the fly the resolution of specific subset of molecules (usually the solvent), which are free to diffuse between the atomistic region and the coarse-grained reservoir. So far, the method has been successfully applied to pure liquids. Coupling the H-AdResS methodology to hybrid models of proteins, such as the Molecular Mechanics/Coarse-Grained (MM/CG) scheme, is a promising approach for rigorous calculations of ligand binding free energies in low-resolution protein models. Towards this goal, here we apply for the first time H-AdResS to two atomistic proteins in dual-resolution solvent, proving its ability to reproduce structural and dynamic properties of both the proteins and the solvent, as obtained from atomistic simulations.

physics.chem-ph

Statistical Analysis of $σ$-Holes: A Novel Complementary View on Halogen Bonding

To contribute to the understanding of noncovalent binding of halogenated molecules with a biological activity, electrostatic potential (ESP) maps of more than 2,500 compounds were thoroughly analysed. A peculiar region of positive ESP, called $σ$-hole, is a concept of central importance for halogen bonding. We aim at simplifying the view on $σ$-holes and provide general trends in organic drug-like molecules. The results are in fair agreement with crystallographic surveys of small molecules as well as of biomolecular complexes and attempt to improve the intuition of chemists when dealing with halogenated compounds.

physics.chem-ph

DNA like$-$charge attraction and overcharging by divalent counterions in the presence of divalent co$-$ions

Strongly correlated electrostatics of DNA systems has drawn the interest of many groups, especially the condensation and overcharging of DNA by multivalent counterions. By adding counterions of different valencies and shapes, one can enhance or reduce DNA overcharging. In this papers, we focus on the effect of multivalent co-ions, specifically divalent co-ions such as SO$_4^{2-}$. A computational experiment of DNA condensation using Monte$-$Carlo simulation in grand canonical ensemble is carried out where DNA system is in equilibrium with a bulk solution containing a mixture of salt of different valency of co-ions. Compared to system with purely monovalent co-ions, the influence of divalent co-ions shows up in multiple aspects. Divalent co-ions lead to an increase of monovalent salt in the DNA condensate. Because monovalent salts mostly participate in linear screening of electrostatic interactions in the system, more monovalent salt molecules enter the condensate leads to screening out of short-range DNA$-$DNA like charge attraction and weaker DNA condensation free energy. The overcharging of DNA by multivalent counterions is also reduced in the presence of divalent co$-$ions. Strong repulsions between DNA and divalent co-ions and among divalent co-ions themselves leads to a {\em depletion} of negative ions near DNA surface as compared to the case without divalent co-ions. At large distance, the DNA$-$DNA repulsive interaction is stronger in the presence of divalent co$-$ions, suggesting that divalent co$-$ions role is not only that of simple stronger linear screening.

q-bio.BM

Proton Dynamics in Protein Mass Spectrometry

Native electrospray ionization/ion mobility-mass spectrometry (ESI/IM-MS) allows an accurate determination of low-resolution structural features of proteins. Yet, the presence of proton dynamics, observed already by us for DNA in the gas phase, and its impact on protein structural determinants, have not been investigated so far. Here, we address this issue by a multi-step simulation strategy on a pharmacologically relevant peptide, the N-terminal residues of amyloid-beta peptide (Abeta(1-16)). Our calculations reproduce the experimental maximum charge state from ESI-MS and are also in fair agreement with collision cross section (CCS) data measured here by ESI/IM-MS. Although the main structural features are preserved, subtle conformational changes do take place in the first ~0.1 ms of dynamics. In addition, intramolecular proton dynamics processes occur on the ps-timescale in the gas phase as emerging from quantum mechanics/molecular mechanics (QM/MM) simulations at the B3LYP level of theory. We conclude that proton transfer phenomena do occur frequently during fly time in ESI-MS experiments (typically on the ms timescale). However, the structural changes associated with the process do not significantly affect the structural determinants.

physics.chem-ph

RNA/peptide binding driven by electrostatics -- Insight from bi-directional pulling simulations

RNA/protein interactions play crucial roles in controlling gene expression. They are becoming important targets for pharmaceutical applications. Due to RNA flexibility and to the strength of electrostatic interactions, standard docking methods are insufficient. We here present a computational method which allows studying the binding of RNA molecules and charged peptides with atomistic, explicit-solvent molecular dynamics. In our method, a suitable estimate of the electrostatic interaction is used as an order parameter (collective variable) which is then accelerated using bi-directional pulling simulations. Since the electrostatic interaction is only used to enhance the sampling, the approximations used to compute it do not affect the final accuracy. The method is employed to characterize the binding of TAR RNA from HIV-1 and a small cyclic peptide. Our simulation protocol allows blindly predicting the binding pocket and pose as well as the binding affinity. The method is general and could be applied to study other electrostatics-driven binding events.

q-bio.BM

Many-Body meets QM/MM: Application to indole in water solution

Spectral properties of chromophores are used to probe complex biological processes in vitro and in vivo, yet how the environment tunes their optical properties is far from being fully understood. Here we present a method to calculate such properties on large scale systems, like biologically relevant molecules in aqueous solution. Our approach is based on many body perturbation theory combined with quantum-mechanics/molecular-mechanics (QM/MM) approach. We show here how to include quasi-particle and excitonic effects for the calculation of optical absorption spectra in a QM/MM scheme. We apply this scheme, together with the well established TDDFT approach, to indole in water solution. Our calculations show that the solvent induces a redshift in the main spectral peak of indole, in quantitative agreement with the experiments and point to the importance of performing averages over molecular dynamics configurations for calculating optical properties.

physics.chem-ph

Convergent dynamics in the protease enzymatic superfamily

Proteases regulate various aspects of the life cycle in all organisms by cleaving specific peptide bonds. Their action is so central for biochemical processes that at least 2% of any known genome encodes for proteolytic enzymes. Here we show that selected proteases pairs, despite differences in oligomeric state, catalytic residues and fold, share a common structural organization of functionally relevant regions which are further shown to undergo similar concerted movements. The structural and dynamical similarities found pervasively across evolutionarily distant clans point to common mechanisms for peptide hydrolysis.

q-bio.BM

Accurate and efficient description of protein vibrational dynamics: comparing molecular dynamics and Gaussian models

Current all-atom potential based molecular dynamics (MD) allow the identification of a protein's functional motions on a wide-range of time-scales, up to few tens of ns. However, functional large scale motions of proteins may occur on a time-scale currently not accessible by all-atom potential based molecular dynamics. To avoid the massive computational effort required by this approach several simplified schemes have been introduced. One of the most satisfactory is the Gaussian Network approach based on the energy expansion in terms of the deviation of the protein backbone from its native configuration. Here we consider an extension of this model which captures in a more realistic way the distribution of native interactions due to the introduction of effective sidechain centroids. Since their location is entirely determined by the protein backbone, the model is amenable to the same exact and computationally efficient treatment as previous simpler models. The ability of the model to describe the correlated motion of protein residues in thermodynamic equilibrium is established through a series of successful comparisons with an extensive (14 ns) MD simulation based on the AMBER potential of HIV-1 protease in complex with a peptide substrate. Thus, the model presented here emerges as a powerful tool to provide preliminary, fast yet accurate characterizations of proteins near-native motion.

cond-mat.stat-mech

Molecular Dynamics Studies on HIV-1 Protease: Drug Resistance and Folding Pathways

Drug resistance to HIV-1 Protease involves accumulation of multiple mutations in the protein. Here we investigate the role of these mutations by using molecular dynamics simulations which exploit the influence of the native-state topology in the folding process. Our calculations show that sites contributing to phenotypic resistance of FDA-approved drugs are among the most sensitive positions for the stability of partially folded states and should play a relevant role in the folding process. Furthermore, associations between amino acid sites mutating under drug treatment are shown to be statistically correlated. The striking correlation between clinical data and our calculations suggest a novel approach to the design of drugs tailored to bind regions crucial not only for protein function but also for folding.

cond-mat.stat-mech

Protein Design is a Key Factor for Subunit-subunit Association

Fundamental questions about the role of the quaternary structures are addressed using a statistical mechanics off-lattice model of a dimer protein. The model, in spite of its simplicity, captures key features of the monomer-monomer interactions revealed by atomic force experiments. Force curves during association and dissociation are characterized by sudden jumps followed by smooth behavior and form hysteresis loops. Furthermore, the process is reversible in a finite range of temperature stabilizing the dimer. It is shown that in the interface between the two monomeric subunits the design procedure naturally favors those amino acids whose mutual interaction is stronger. Furthermore it is shown that the width of the hysteresis loop increases as the design procedure improves, i.e. stabilizes more the dimer.

cond-mat.stat-mech