SearcharxivSearch

arXiv subjects

Daniel M. Zuckerman

Publications and source records attributed to Daniel M. Zuckerman.

At least 19 recordsLinked to original sources

Markov state models revisited: Principles and algorithms for unbiased observables

Markov state models (MSMs) have become ubiquitous tools for analyzing molecular dynamics (MD) simulations because of their simple, powerful premise: although complete MD sampling may be impossible, the MSM can "stitch together" transition probabilities derived from local sampling to provide a global picture of kinetics and mechanisms. In the standard MSM framework, the available MD data is organized into a single transition matrix, which is then used to estimate all observables at a lag time chosen so the coarse-grained dynamics are approximately Markovian. This approach leads to avoidable model bias and motivates long lag times that obscure short-timescale processes of interest. In contrast, this paper shows how to obtain unbiased coarse-grained observables at any fixed lag time and for any fixed coarse-graining in the limit of infinite, properly weighted data. The central idea is to replace the single-matrix framework with two transition matrices -- one representing equilibrium dynamics and another representing source-sink recycling dynamics -- and use the correct matrix or matrices to estimate the matched dynamical observables.

cond-mat.stat-mech

Reducing Weighted Ensemble Variance With Optimal Trajectory Management

Weighted ensemble (WE) is an enhanced path-sampling method that is conceptually simple, widely applicable, and statistically exact. In a WE simulation, an ensemble of trajectories is periodically pruned or replicated to enhance sampling of rare transitions and improve estimation of mean first passage times (MFPTs). However, poor choices of the parameters governing pruning and replication can lead to high-variance MFPT estimates. Our previous work [J. Chem. Phys. 158, 014108 (2023)] presented an optimal WE parameterization strategy and applied it in low-dimensional example systems. The strategy harnesses estimated local MFPTs from different initial configurations to a single target state. In the present work, we apply the optimal parameterization strategy to more challenging, high-dimensional molecular models, namely, synthetic molecular dynamics (MD) models of Trp-cage folding and unfolding, as well as atomistic MD models of NTL9 folding in high-friction and low-friction continuum solvents. In each system we use WE to estimate the MFPT for folding or unfolding events. We show that the optimal parameterization reduces the variance of MFPT estimates in three of four systems, with dramatic improvement in the most challenging atomistic system. Overall, the parameterization strategy improves the accuracy and reliability of WE estimates for the kinetics of biophysical processes.

physics.chem-ph

RiteWeight: Randomized Iterative Trajectory Reweighting for Steady-State Distributions Without Discretization Error

A significant challenge in molecular dynamics (MD) simulations is ensuring that sampled configurations converge to the equilibrium or nonequilibrium stationary distribution of interest. Lack of convergence constrains the estimation of free energies, rates, and mechanisms of complex molecular events. Here, we introduce the "Randomized ITErative trajectory reWeighting" (RiteWeight) algorithm to estimate a stationary distribution from unconverged simulation data. This method iteratively reweights trajectory segments in a self-consistent way by solving for the stationary distribution of a Markov state model (MSM), updating segment weights, and employing a new random clustering in each iteration. The iterative random clustering mitigates the phase-space discretization error inherent in existing trajectory reweighting techniques and yields quasi-continuous configuration-space distributions. We present mathematical analysis of the algorithm's fixed points as well as empirical validation using both synthetic MD Trp-cage trajectories, for which the stationary solution is exactly calculable, and standard atomistic MD Trp-cage trajectories extracted from a long reference simulation. In both test systems, we find that RiteWeight corrects flawed distributions and generates accurate observables for equilibrium and nonequilibrium steady states. The results highlight the value of correcting the underlying trajectory distribution rather than using a standard MSM

physics.comp-ph

Simple synthetic molecular dynamics for efficient trajectory generation

Synthetic molecular dynamics (synMD) trajectories from learned generative models have been proposed as a useful addition to the biomolecular simulation toolbox. The computational expense of explicitly integrating the equations of motion in molecular dynamics currently is a severe limit on the number and length of trajectories which can be generated for complex systems. Approximate, but more computationally efficient, generative models can be used in place of explicit integration of the equations of motion, and can produce meaningful trajectories at greatly reduced computational cost. Here, we demonstrate a very simple synMD approach using a fine-grained Markov state model (MSM) with states mapped to specific atomistic configurations, which provides an exactly solvable reference. We anticipate this simple approach will enable rapid, effective testing of enhanced sampling algorithms in highly non-trivial models for both equilibrium and non-equilibrium problems. We demonstrate the use of a MSM to generate atomistic synMD trajectories for the fast-folding miniprotein Trp-cage, at a rate of over 200 milliseconds per day on a standard workstation. We employ a non-standard clustering for MSM generation that appears to better preserve kinetic properties at shorter lag times than a conventional MSM. We also show a parallelizable workflow that backmaps discrete synMD trajectories to full-coordinate representations at dynamic resolution for efficient analysis.

physics.comp-ph

A gentle introduction to the non-equilibrium physics of trajectories: Theory, algorithms, and biomolecular applications

Despite the importance of non-equilibrium statistical mechanics in modern physics and related fields, the topic is often omitted from undergraduate and core-graduate curricula. Key aspects of non-equilibrium physics, however, can be understood with a minimum of formalism based on a rigorous trajectory picture. The fundamental object is the ensemble of trajectories, a set of independent time-evolving systems that easily can be visualized or simulated (for protein folding, e.g.), and which can be analyzed rigorously in analogy to an ensemble of static system configurations. The trajectory picture provides a straightforward basis for understanding first-passage times, "mechanisms" in complex systems, and fundamental constraints the apparent reversibility of complex processes. Trajectories make concrete the physics underlying the diffusion and Fokker-Planck partial differential equations. Last but not least, trajectory ensembles underpin some of the most important algorithms which have provided significant advances in biomolecular studies of protein conformational and binding processes.

cond-mat.stat-mech

Unbiased estimation of equilibrium, rates, and committors from Markov state model analysis

Markov state models (MSMs) have been broadly adopted for analyzing molecular dynamics trajectories, but the approximate nature of the models that results from coarse-graining into discrete states is a long-known limitation. We show theoretically that, despite the coarse graining, in principle MSM-like analysis can yield unbiased estimation of key observables. We describe unbiased estimators for equilibrium state populations, for the mean first-passage time (MFPT) of an arbitrary process, and for state committors - i.e., splitting probabilities. Generically, the estimators are only asymptotically unbiased but we describe how extension of a recently proposed reweighting scheme can accelerate relaxation to unbiased values. Exactly accounting for 'sliding window' averaging over finite-length trajectories is a key, novel element of our analysis. In general, our analysis indicates that coarse-grained MSMs are asymptotically unbiased for steady-state properties only when appropriate boundary conditions (e.g., source-sink for MFPT estimation) are applied directly to trajectories, prior to calculation of the appropriate transition matrix.

physics.comp-ph

Iterative trajectory reweighting for estimation of equilibrium and non-equilibrium observables

We present two algorithms by which a set of short, unbiased trajectories can be iteratively reweighted to obtain various observables. The first algorithm estimates the stationary (steady state) distribution of a system by iteratively reweighting the trajectories based on the average probability in each state. The algorithm applies to equilibrium or non-equilibrium steady states, exploiting the `left' stationarity of the distribution under dynamics -- i.e., in a discrete setting, when the column vector of probabilities is multiplied by the transition matrix expressed as a left stochastic matrix. The second procedure relies on the `right' stationarity of the committor (splitting probability) expressed as a row vector. The algorithms are unbiased, do not rely on computing transition matrices, and make no Markov assumption about discretized states. Here, we apply the procedures to a one-dimensional double-well potential, and to a 208$μ$s atomistic Trp-cage folding trajectory from D.E. Shaw Research.

physics.comp-ph

Optimizing weighted ensemble sampling of steady states

We propose parameter optimization techniques for weighted ensemble sampling of Markov chains in the steady-state regime. Weighted ensemble consists of replicas of a Markov chain, each carrying a weight, that are periodically resampled according to their weights inside of each of a number of bins that partition state space. We derive, from first principles, strategies for optimizing the choices of weighted ensemble parameters, in particular the choice of bins and the number of replicas to maintain in each bin. In a simple numerical example, we compare our new strategies with more traditional ones and with direct Monte Carlo.

math.NA

Key biology you should have learned in physics class: Using ideal-gas mixtures to understand biomolecular machines

The biological cell exhibits a fantastic range of behaviors, but ultimately these are governed by a handful of physical and chemical principles. Here we explore simple theory, known for decades and based on the simple thermodynamics of mixtures of ideal gases, which illuminates several key functions performed within the cell. Our focus is the free-energy-driven import and export of molecules, such as nutrients and other vital compounds, via transporter proteins. Complementary to a thermodynamic picture is a description of transporters via "mass-action" chemical kinetics, which lends further insights into biological machinery and free energy use. Both thermodynamic and kinetic descriptions can shed light on the fundamental non-equilibrium aspects of transport. On the whole, our biochemical-physics discussion will remain agnostic to chemical details, but we will see how such details ultimately enter a physical description through the example of the cellular fuel ATP.

q-bio.BM

Transient probability currents provide upper and lower bounds on non-equilibrium steady-state currents in the Smoluchowski picture

Probability currents are fundamental in characterizing the kinetics of non-equilibrium processes. Notably, the steady-state current $J_{ss}$ for a source-sink system can provide the exact mean-first-passage time (MFPT) for the transition from source to sink. Because transient non-equilibrium behavior is quantified in some modern path sampling approaches, such as the "weighted ensemble" strategy, there is strong motivation to determine bounds on $J_{ss}$ -- and hence on the MFPT -- as the system evolves in time. Here we show that $J_{ss}$ is bounded from above and below by the maximum and minimum, respectively, of the current as a function of the spatial coordinate at any time $t$ for one-dimensional systems undergoing over-damped Langevin (i.e., Smoluchowski) dynamics and for higher-dimensional Smoluchowski systems satisfying certain assumptions when projected onto a single dimension. These bounds become tighter with time, making them of potential practical utility in a scheme for estimating $J_{ss}$ and the long-timescale kinetics of complex systems. Conceptually, the bounds result from the fact that extrema of the transient currents relax toward the steady-state current.

cond-mat.stat-mech

Statistical uncertainty analysis for small-sample, high log-variance data: Cautions for bootstrapping and Bayesian bootstrapping

Recent advances in molecular simulations allow the evaluation of previously unattainable observables, such as rate constants for protein folding. However, these calculations are usually computationally expensive and even significant computing resources may result in a small number of independent estimates spread over many orders of magnitude. Such small-sample, high "log-variance" data are not readily amenable to analysis using the standard uncertainty (i.e., "standard error of the mean") because unphysical negative limits of confidence intervals result. Bootstrapping, a natural alternative guaranteed to yield a confidence interval within the minimum and maximum values, also exhibits a striking systematic bias of the lower confidence limit in log space. As we show, bootstrapping artifactually assigns high probability to improbably low mean values. A second alternative, the Bayesian bootstrap strategy, does not suffer from the same deficit and is more logically consistent with the type of confidence interval desired. The Bayesian bootstrap provides uncertainty intervals that are more reliable than those from the standard bootstrap method, but must be used with caution nevertheless. Neither standard nor Bayesian bootstrapping can overcome the intrinsic challenge of under-estimating the mean from small-size, high log-variance samples. Our conclusions are based on extensive analysis of model distributions and re-analysis of multiple independent atomistic simulations. Although we only analyze rate constants, similar considerations will apply to related calculations, potentially including highly non-linear averages like the Jarzynski relation.

stat.AP

Pathway Histogram Analysis of Trajectories: A general strategy for quantification of molecular mechanisms

A key overall goal of biomolecular simulations is the characterization of "mechanism" -- the pathways through configuration space of processes such as conformational transitions and binding. Some amount of heterogeneity is intrinsic to the ensemble of pathways, in direct analogy to thermal configurational ensembles. Quantification of that heterogeneity is essential to a complete understanding of mechanism. We propose a general approach for characterizing path ensembles based on mapping individual trajectories into pathway classes whose populations and uncertainties can be analyzed as an ordinary histogram, providing a quantitative "fingerprint" of mechanism. In contrast to prior flux-based analyses used for discrete-state models, stochastic deviations from average behavior are explicitly included via direct classification of trajectories. The histogram approach, furthermore, is applicable to analysis of continuous trajectories. It enables straightforward comparison between ensembles produced by different methods or under different conditions. To implement the formulation, we develop approaches for classifying trajectories, including a clustering-based approach suitable for both continuous-space (e.g., molecular dynamics) or discrete-state (e.g., Markov state model) trajectories, as well as a "fundamental sequence" approach tailored for discrete-state trajectories but also applicable to continuous trajectories through a mapping process. We apply the pathway histogram analysis to a toy model and an extremely long atomistic molecular dynamics trajectory of protein folding.

physics.chem-ph

Stochastic Simulation to Visualize Gene Expression and Error Correction in Living Cells

Stochastic simulation can make the molecular processes of cellular control more vivid than the traditional differential-equation approach by generating typical system histories instead of just statistical measures such as the mean and variance of a population. Simple simulations are now easy for students to construct from scratch, that is, without recourse to black-box packages. In some cases, their results can also be compared directly to single-molecule experimental data. After introducing the stochastic simulation algorithm, this article gives two case studies, involving gene expression and error correction, respectively. Code samples and resulting animations showing results are given in the online supplements.

q-bio.SC

Biophysical comparison of ATP-driven proton pumping mechanisms suggests a kinetic advantage for the rotary process depending on coupling ratio

ATP-driven proton pumps, which are critical to the operation of a cell, maintain cytosolic and organellar pH levels within a narrow functional range. These pumps employ two very different mechanisms: an elaborate rotary mechanism used by V-ATPase H+ pumps, and a simpler alternating access mechanism used by P-ATPase H+ pumps. Why are two different mechanisms used to perform the same function? Systematic analysis, without parameter fitting, of kinetic models of the rotary, alternating access and other possible mechanisms suggest that, when the ratio of protons transported per ATP hydrolyzed exceeds one, the one-at-a-time proton transport by the rotary mechanism is faster than other possible mechanisms across a wide range of driving conditions. When the ratio is one, there is no intrinsic difference in the free energy landscape between mechanisms, and therefore all mechanisms can exhibit the same kinetic performance. To our knowledge all known rotary pumps have an H+:ATP ratio greater than one, and all known alternating access ATP-driven proton pumps have a ratio of one. Our analysis suggests a possible explanation for this apparent relationship between coupling ratio and mechanism. When the conditions under which the pump must operate permit a coupling ratio greater than one, the rotary mechanism may have been selected for its kinetic advantage. On the other hand, when conditions require a coupling ratio of one or less, the alternating access mechanism may have been selected for other possible advantages resulting from its structural and functional simplicity.

q-bio.SC

A proposal for regularly updated review/survey articles: "Perpetual Reviews"

We advocate the publication of review/survey articles that will be updated regularly, both in traditional journals and novel venues. We call these "perpetual reviews." This idea naturally builds on the dissemination and archival capabilities present in the modern internet, and indeed perpetual reviews exist already in some forms. Perpetual review articles allow authors to maintain over time the relevance of non-research scholarship that requires a significant investment of effort. Further, such reviews published in a purely electronic format without space constraints can also permit more pedagogical scholarship and clearer treatment of technical issues that remain obscure in a brief treatment.

cs.DL

Simultaneous computation of dynamical and equilibrium information using a weighted ensemble of trajectories

Equilibrium formally can be represented as an ensemble of uncoupled systems undergoing unbiased dynamics in which detailed balance is maintained. Many non-equilibrium processes can be described by suitable subsets of the equilibrium ensemble. Here, we employ the "weighted ensemble" (WE) simulation protocol [Huber and Kim, Biophys. J., 1996] to generate equilibrium trajectory ensembles and extract non-equilibrium subsets for computing kinetic quantities. States do not need to be chosen in advance. The procedure formally allows estimation of kinetic rates between arbitrary states chosen after the simulation, along with their equilibrium populations. We also describe a related history-dependent matrix procedure for estimating equilibrium and non-equilibrium observables when phase space has been divided into arbitrary non-Markovian regions, whether in WE or ordinary simulation. In this proof-of-principle study, these methods are successfully applied and validated on two molecular systems: explicitly solvated methane association and the implicitly solvated Ala4 peptide. We comment on challenges remaining in WE calculations.

physics.bio-ph

Efficient Stochastic Simulation of Chemical Kinetics Networks using a Weighted Ensemble of Trajectories

We apply the "weighted ensemble" (WE) simulation strategy, previously employed in the context of molecular dynamics simulations, to a series of systems-biology models that range in complexity from one-dimensional to a system with 354 species and 3680 reactions. WE is relatively easy to implement, does not require extensive hand-tuning of parameters, does not depend on the details of the simulation algorithm, and can facilitate the simulation of extremely rare events. For the coupled stochastic reaction systems we study, WE is able to produce accurate and efficient approximations of the joint probability distribution for all chemical species for all time t. WE is also able to efficiently extract mean first passage times for the systems, via the construction of a steady-state condition with feedback. In all cases studied here, WE results agree with independent calculations, but significantly enhance the precision with which rare or slow processes can be characterized. Speedups over "brute-force" in sampling rare events via the Gillespie direct Stochastic Simulation Algorithm range from ~10^12 to ~10^20 for rare states in a distribution, and ~10^2 to ~10^4 for finding mean first passage times.

q-bio.MN

Equilibrium Sampling in Biomolecular Simulation

Equilibrium sampling of biomolecules remains an unmet challenge after more than 30 years of atomistic simulation. Efforts to enhance sampling capability, which are reviewed here, range from the development of new algorithms to parallelization to novel uses of hardware. Special focus is placed on classifying algorithms -- most of which are underpinned by a few key ideas -- in order to understand their fundamental strengths and limitations. Although algorithms have proliferated, progress resulting from novel hardware use appears to be more clear-cut than from algorithms alone, partly due to the lack of widely used sampling measures.

q-bio.BM