SearcharxivSearch

arXiv subjects

Gerhard Hummer

Publications and source records attributed to Gerhard Hummer.

At least 19 recordsLinked to original sources

Bayesian Learning of Distance Metrics Beyond RMSD for Biomolecule Alignment, Clustering, and Domain Identification

Root-mean-square deviation (RMSD) is the standard metric of structural comparison in molecular dynamics (MD) simulations. In its conventional form, RMSD assigns equal weight to all atoms regardless of mobility. Hence, flexible loops and disordered regions can dominate a global RMSD, while the rigid functional core contributes negligibly to the overall metric. To address this issue, we introduce the Bayes-optimal RMSD (BRMSD), which optimizes per-atom weights jointly with structural averages by maximizing a Bayesian posterior. In a trade-off between low RMSD and weight uniformity, a position-fluctuation parameter $\sigma$ controls the transition from classical RMSD ($\sigma \to \infty$) to a progressive focus on a rigid core ($\sigma \to 0$). The BRMSD framework supports analysis modules for structural alignment, focused alignment onto a user-specified domain, trajectory smoothing, soft $K$-means conformational clustering, and rigid-domain identification. These modules are implemented in the open-source Python package BRMSD (https://github.com/bio-phys/BRMSD) and benchmarked on two MD systems, the endoplasmic reticulum translocon-associated protein SND3 and the phosphotransferase adenylate kinase.

q-bio.BM

Extrapolating molecular dynamics simulations to zero time step and across thermodynamic space

The integration time step is a critical determinant of performance in molecular dynamics simulations, governing the trade-off between speed and fidelity. Although 2 fs remains the standard in atomistic biomolecular simulations, the push for performance has popularized a 4 fs time step with hydrogen mass repartitioning, often combined with multiple time stepping or mass rescaling. However, it is often unclear whether a chosen protocol is overly aggressive, as the apparent numerical stability of a trajectory can mask underlying thermodynamic inaccuracies. Increasing the time step will exacerbate systematic discretization errors, inherent to all numerical integration algorithms. In the widely used Verlet family of integrators, these errors manifest as $\mathcal{O}(\Delta t^2)$ deviations in thermodynamic observables such as potential energy and volume, and for common Langevin splitting schemes, even temperature. We demonstrate that these deviations follow a simple, linear thermodynamic model, allowing for their rigorous removal by extrapolation to the zero time step limit. In turn, the time-step dependence provides us with estimates of the system heat capacity, compressibility, and thermal expansion coefficient. This framework allows us to construct consistent probability distributions of energy and volume across thermodynamic states, effectively recovering Boltzmann-consistent statistics at a target condition independent of time step. These considerations are particularly important for enhanced sampling methods such as replica exchange and umbrella sampling, which rely on rigorous Boltzmann sampling and require accurate energies and temperatures for valid replica exchange probabilities and statistical reweighting.

physics.bio-ph

Random functions as data compressors for machine learning of molecular processes

Machine learning (ML) is rapidly transforming the way molecular dynamics simulations are performed and analyzed, from materials modeling to studies of protein folding and function. ML algorithms are often employed to learn low-dimensional representations of conformational landscapes and to cluster trajectories into relevant metastable states. Most of these algorithms require selecting a small number of features that describe the problem of interest. Although deep neural networks can tackle large numbers of input features, the training costs increase with input size, which makes the selection of a subset of features mandatory for most problems of practical interest. Here, we show that random nonlinear projections can be used to compress large feature spaces and make computations faster without substantial loss of information. We describe an efficient way to produce random projections and then exemplify the general procedure for protein folding. For our test cases NTL9 and the double-norleucin variant of the villin headpiece, we find that random compression retains the core static and dynamic information of the original high dimensional feature space and makes trajectory analysis more robust.

cond-mat.soft

AI-guided transition path sampling of lipid flip-flop and membrane nanoporation

We study lipid translocation ("flip-flop") between the leaflets of a planar lipid bilayer with transition path sampling (TPS). Rare flip-flops compete with biological machineries that actively establish asymmetric lipid compositions. Artificial Intelligence (AI) guided TPS captures flip-flop without biasing the dynamics by initializing molecular dynamics simulations close to the tipping point, i.e., where it is equally likely for a lipid to next go to one or the other leaflet. We train a neural network model on the fly to predict the respective probability, i.e., the "committor" encoding the mechanism of flip-flop. Whereas coarse-grained DMPC lipids "tunnel" through the hydrophobic bilayer, unaided by water, atomistic DMPC lipids instead utilize spontaneously formed water nanopores to traverse to the other side. For longer DSPC lipids, these membrane defects are less stable, with lipid transfer along transient water threads in a locally thinned membrane emerging as a third distinct mechanism. Remarkably, in the high (~660) dimensional feature space of the deep neural networks, the reaction coordinate becomes effectively linear, in line with Cover's theorem and consistent with the idea of dominant reaction tubes.

cond-mat.soft

The need to implement FAIR principles in biomolecular simulations

This letter illustrates the opinion of the molecular dynamics (MD) community on the need to adopt a new FAIR paradigm for the use of molecular simulations. It highlights the necessity of a collaborative effort to create, establish, and sustain a database that allows findability, accessibility, interoperability, and reusability of molecular dynamics simulation data. Such a development would democratize the field and significantly improve the impact of MD simulations on life science research. This will transform our working paradigm, pushing the field to a new frontier. We invite you to support our initiative at the MDDB community (https://mddbr.eu/community/) Now published as: Amaro, R.E., et al. The need to implement FAIR principles in biomolecular simulations. Nat Methods (2025) https://doi.org/10.1038/s41592-025-02635-0

q-bio.BM

Nanosecond chain dynamics of single-stranded nucleic acids

The conformational dynamics of single-stranded nucleic acids are fundamental for nucleic acid folding and function. However, their elementary chain dynamics have been difficult to resolve experimentally. Here we employ a combination of single-molecule F\"orster resonance energy transfer, nanosecond fluorescence correlation spectroscopy, fluorescence lifetime analysis, and nanophotonic enhancement to determine the conformational ensembles and rapid chain dynamics of short single-stranded nucleic acids in solution. To interpret the experimental results in terms of end-to-end distance dynamics, we utilize the hierarchical chain growth approach, simple polymer models, and refinement with Bayesian inference of ensembles to generate structural ensembles that closely align with the experimental data. The resulting chain reconfiguration times are exceedingly rapid, in the 10-ns range. Solvent viscosity-dependent measurements indicate that these dynamics of single-stranded nucleic acids exhibit negligible internal friction and are thus dominated by solvent friction. Our results provide a detailed view of the conformational distributions and rapid dynamics of single-stranded nucleic acids.

physics.bio-ph

Unwrapping NPT simulations to calculate diffusion coefficients

In molecular dynamics simulations in the NPT ensemble at constant pressure, the size and shape of the periodic simulation box fluctuate with time. For particle images far from the origin, the rescaling of the box by the barostat results in unbounded position displacements. Special care is thus required when a particle trajectory is unwrapped from a projection into the central box under periodic boundary conditions to a trajectory in full three-dimensional space, e.g., for the calculation of diffusion coefficients. Here, we review and compare different schemes in use for trajectory unwrapping. We also specify the corresponding rewrapping schemes to put an unwrapped trajectory back into the central box. On this basis, we then identify a scheme for the calculation of meaningful diffusion coefficients, which is a primary application of trajectory unwrapping. In this scheme, the wrapped and unwrapped trajectory are mutually consistent and their statistical properties are preserved. We conclude with advice on best practice for the consistent unwrapping of constant-pressure simulation trajectories and the calculation of accurate translational diffusion coefficients.

physics.comp-ph

Efficient generation of random rotation matrices in four dimensions

Markov-chain Monte Carlo algorithms rely on trial moves that are either rejected or accepted based on certain criteria. Here, we provide an efficient algorithm to generate random rotation matrices in four dimensions (4D) covering an arbitrary pre-defined range of rotation angles. The matrices can be combined with Monte Carlo methods for the efficient sampling of the SO(4) group of 4D rotations. The matrices are unbiased and constructed such that repeated rotations result in uniform sampling over SO(4). 4D rotations can be used to optimize the mass partitioning for stable time integration in coarse-grained molecular dynamics simulations and should find further applications in the fields of robotics and computer vision.

physics.comp-ph

Structural ensembles of disordered proteins from hierarchical chain growth and simulation

Disordered proteins and nucleic acids play key roles in cellular function and disease. Here we review recent advances in the computational exploration of the conformational dynamics of flexible biomolecules. We focus on hierarchical chain growth (HCG) from fragment libraries built with atomistic molecular dynamics simulations. HCG combines chain fragments in a statistically reproducible manner into ensembles of full-length atomically detailed biomolecular structures. The input fragment structures are typically collected from molecular dynamics simulations, but could also come from structural databases. Experimental data can be integrated during and after chain assembly. Applications to the neurodegeneration-linked proteins $\alpha$-synuclein, tau, and TDP-43, including as condensate, illustrate the use of HCG. We conclude by highlighting the emerging connections to AI-based structural modeling.

physics.chem-ph

Rebinding kinetics from single-molecule force spectroscopy experiments close to equilibrium

Analysis of bond rupture data from single-molecule force spectroscopy experiments commonly relies on the strong assumption that the bond dissociation process is irreversible. However, with increased spatiotemporal resolution of instruments it is now possible to observe multiple unbinding-rebinding events in a single pulling experiment. Here, we augment the theory of force-induced unbinding by explicitly taking into account rebinding kinetics, and provide approximate analytic solutions of the resulting rate equations. Furthermore, we use a short-time expansion of the exact kinetics to construct numerically efficient maximum likelihood estimators for the parameters of the force-dependent unbinding and rebinding rates, which pair well with and complement established methods, such as the analysis of rate maps. We provide an open-source implementation of the theory, evaluated for Bell-like rates, which we apply to synthetic data generated by a Gillespie stochastic simulation algorithm for time-dependent rates.

physics.bio-ph

Small ionic radii limit time step in Martini 3 molecular dynamics simulations

Among other improvements, the Martini 3 coarse-grained force field provides a more accurate description of the solvation of protein pockets and channels through the consistent use of various bead types and sizes. Here, we show that the representation of Na$^+$ and Cl$^-$ ions as "tiny" (TQ5) beads limits the accessible time step to 25 fs. By contrast, with Martini 2, time steps of 30-40 fs were possible for lipid bilayer systems without proteins. This limitation is relevant for, e.g., phase separating lipid mixtures that require long equilibration times. We derive a quantitative kinetic model of time-integration instabilities in molecular dynamics (MD) as a function of time step, ion concentration and mass, system size, and simulation time. With this model, we demonstrate that ion-water interactions are the main source of instability at physiological conditions, followed closely by ion-ion interactions. We show that increasing the ionic masses makes it possible to use time steps up to 40 fs with minimal impact on static equilibrium properties and on dynamical quantities such as lipid and ion diffusion coefficients. Increasing the size of the bead representing the ions (and thus changing their hydration) also permits longer time steps. The use of larger time steps in Martini 3 simulations results in a more efficient exploration of configuration space. The kinetic model of MD simulation crashes can be used to determine the maximum allowed time step whenever sampling efficiency is critical.

cond-mat.soft

Transition rates, survival probabilities, and quality of bias from time-dependent biased simulations

Simulations with an adaptive time-dependent bias, such as metadynamics, enable an efficient exploration of the conformational space of a system. However, the dynamic information of the system is altered by the bias. With infrequent metadynamics it is possible to recover the transition rate of crossing a barrier, if the collective variables are ideal and there is no bias deposition near the transition state. Unfortunately, for simulations of complex molecules, these conditions are not always fulfilled. To overcome these limitations, and inspired by single-molecule force spectroscopy, we developed a method based on Kramers' theory for calculating the barrier-crossing rate when a time-dependent bias is added to the system. We assess the quality of the bias parameter by measuring how efficiently the bias accelerates the transitions compared to ideal behavior. We present approximate analytical expressions of the survival probability that accurately reproduce the barrier-crossing time statistics, and enable the extraction of the unbiased transition rate even for challenging cases, where previous methods fail.

physics.bio-ph

Autonomous artificial intelligence discovers mechanisms of molecular self-organization in virtual experiments

Molecular self-organization driven by concerted many-body interactions produces the ordered structures that define both inanimate and living matter. Understanding the physical mechanisms that govern the formation of molecular complexes and crystals is key to controlling the assembly of nanomachines and new materials. We present an artificial intelligence (AI) agent that uses deep reinforcement learning and transition path theory to discover the mechanism of molecular self-organization phenomena from computer simulations. The agent adaptively learns how to sample complex molecular events and, on the fly, constructs quantitative mechanistic models. By using the mechanistic understanding for AI-driven sampling, the agent closes the learning cycle and overcomes time-scale gaps of many orders of magnitude. Symbolic regression condenses the mechanism into a human-interpretable form. Applied to ion association in solution, gas-hydrate crystal formation, and membrane-protein assembly, the AI agent identifies the many-body solvent motions governing the assembly process, discovers the variables of classical nucleation theory, and reveals competing assembly pathways. The mechanistic descriptions produced by the agent are predictive and transferable to close thermodynamic states and similar systems. Autonomous AI sampling has the power to discover assembly and reaction mechanisms from materials science to biology.

physics.chem-ph

Maximum likelihood estimates of diffusion coefficients from single-particle tracking experiments

Single-molecule localization microscopy allows practitioners to locate and track labeled molecules in biological systems. When extracting diffusion coefficients from the resulting trajectories, it is common practice to perform a linear fit on mean-square-displacement curves. However, this strategy is suboptimal and prone to errors. Recently, it was shown that the increments between observed positions provide a good estimate for the diffusion coefficient, and their statistics are well-suited for likelihood-based analysis methods. Here, we revisit the problem of extracting diffusion coefficients from single-particle tracking experiments subject to static and dynamic noise using the principle of maximum likelihood. Taking advantage of an efficient real-space formulation, we extend the model to mixtures of subpopulations differing in their diffusion coefficients, which we estimate with the help of the expectation-maximization algorithm. This formulation naturally leads to a probabilistic assignment of trajectories to subpopulations. We employ the theory to analyze experimental tracking data that cannot be explained with a single diffusion coefficient. We test how well a dataset conforms to the assumptions of a diffusion model and determine the optimal number of subpopulations with the help of a quality factor of known analytical distribution. To facilitate use by practitioners, we provide a fast open-source implementation of the theory for the efficient analysis of multiple trajectories in arbitrary dimensions simultaneously.

physics.bio-ph

Optimal estimates of diffusion coefficients from molecular dynamics simulations

Translational diffusion coefficients are routinely estimated from molecular dynamics simulations. Linear fits to mean squared displacement (MSD) curves have become the de facto standard, from simple liquids to complex biomacromolecules. Nonlinearities in MSD curves at short times are handled with a wide variety of ad hoc practices, such as partial and piece-wise fitting of the data. Here, we present a rigorous framework to obtain reliable estimates of the diffusion coefficient and its statistical uncertainty. We also assess in a quantitative manner if the observed dynamics is indeed diffusive. By accounting for correlations between MSD values at different times, we reduce the statistical uncertainty of the estimator and thereby increase its efficiency. With a Kolmogorov-Smirnov test, we check for possible anomalous diffusion. We provide an easy-to-use Python data analysis script for the estimation of diffusion coefficients. As an illustration, we apply the formalism to molecular dynamics simulation data of pure TIP4P-D water and a single ubiquitin protein. In a companion paper [J. Chem. Phys. XXX, YYYYY (2020)], we demonstrate its ability to recognize deviations from regular diffusion caused by systematic errors in a common trajectory "unwrapping" scheme that is implemented in popular simulation and visualization software.

physics.comp-ph

Systematic errors in diffusion coefficients from long-time molecular dynamics simulations at constant pressure

In molecular dynamics simulations under periodic boundary conditions, particle positions are typically wrapped into a reference box. For diffusion coefficient calculations using the Einstein relation, the particle positions need to be unwrapped. Here, we show that a widely used heuristic unwrapping scheme is not suitable for long simulations at constant pressure. Improper accounting for box-volume fluctuations creates, at long times, unphysical trajectories and, in turn, grossly exaggerated diffusion coefficients. We propose an alternative unwrapping scheme that resolves this issue. At each time step, we add the minimal displacement vector according to periodic boundary conditions for the instantaneous box geometry. Here and in a companion paper [J. Chem. Phys. XXX, YYYYY (2020)], we apply the new unwrapping scheme to extensive molecular dynamics and Brownian dynamics simulation data. We provide practitioners with a formula to assess if and by how much earlier results might have been affected by the widely used heuristic unwrapping scheme.

physics.comp-ph

Cross-validation tests for cryo-EM maps using an independent particle set

Cryo-electron microscopy is a revolutionary technique that can provide 3D density maps at near-atomic resolution. However, map validation is still an open issue in the field. Despite several efforts from the community, it is possible to overfit the reconstructions to noisy data. Here, inspired by modern statistics, we develop a novel methodology that uses a small independent particle set to validate the 3D maps. The main idea is to monitor how the map probability evolves over the control set during the refinement. The method is complementary to the gold-standard procedure, which generates two reconstructions at each iteration. We low-pass filter the two reconstructions for different frequency cutoffs, and we calculate the probability of each filtered map given the control set. For high-quality maps, the probability should increase as a function of the frequency cutoff and of the refinement iteration. We also compute the similarity between the probability distributions of the two reconstructions. As higher frequencies are added to the maps, more dissimilar are the distributions. We optimized the BioEM software package to perform these calculations, and tested the method on several systems, some which were overfitted. Our results show that our method is able to discriminate the overfitted sets from the non-overfitted ones. We conclude that having a control particle set, not used for the refinement, is essential for cross-validating cryo-EM maps.

physics.bio-ph

Molecular free energy profiles from force spectroscopy experiments by inversion of observed committors

In single-molecule force spectroscopy experiments, a biomolecule is attached to a force probe via polymer linkers, and the total extension -- of molecule plus apparatus -- is monitored as a function of time. In a typical unfolding experiment at constant force, the total extension jumps between two values that correspond to the folded and unfolded states of the molecule. For several biomolecular systems the committor, which is the probability to fold starting from a given extension, has been used to extract the molecular activation barrier (a technique known as "committor inversion"). In this work, we study the influence of the force probe, which is much larger than the molecule being measured, on the activation barrier obtained by committor inversion. We use a two-dimensional framework in which the diffusion coefficient of the molecule and of the pulling device can differ. We systematically study the free energy profile along the total extension obtained from the committor, by numerically solving the Onsager equation and using Brownian dynamics simulations. We analyze the dependence of the extracted barrier on the linker stiffness, molecular barrier height, and diffusion anisotropy, and thus, establish the range of validity of committor inversion. Along the way, we showcase the committor of 2-dimensional diffusive models and illustrate how it is affected by barrier asymmetry and diffusion anisotropy.

physics.chem-ph