SearcharxivSearch

arXiv subjects

David Aristoff

Publications and source records attributed to David Aristoff.

At least 19 recordsLinked to original sources

Markov state models revisited: Principles and algorithms for unbiased observables

Markov state models (MSMs) have become ubiquitous tools for analyzing molecular dynamics (MD) simulations because of their simple, powerful premise: although complete MD sampling may be impossible, the MSM can "stitch together" transition probabilities derived from local sampling to provide a global picture of kinetics and mechanisms. In the standard MSM framework, the available MD data is organized into a single transition matrix, which is then used to estimate all observables at a lag time chosen so the coarse-grained dynamics are approximately Markovian. This approach leads to avoidable model bias and motivates long lag times that obscure short-timescale processes of interest. In contrast, this paper shows how to obtain unbiased coarse-grained observables at any fixed lag time and for any fixed coarse-graining in the limit of infinite, properly weighted data. The central idea is to replace the single-matrix framework with two transition matrices -- one representing equilibrium dynamics and another representing source-sink recycling dynamics -- and use the correct matrix or matrices to estimate the matched dynamical observables.

cond-mat.stat-mech

Reducing Weighted Ensemble Variance With Optimal Trajectory Management

Weighted ensemble (WE) is an enhanced path-sampling method that is conceptually simple, widely applicable, and statistically exact. In a WE simulation, an ensemble of trajectories is periodically pruned or replicated to enhance sampling of rare transitions and improve estimation of mean first passage times (MFPTs). However, poor choices of the parameters governing pruning and replication can lead to high-variance MFPT estimates. Our previous work [J. Chem. Phys. 158, 014108 (2023)] presented an optimal WE parameterization strategy and applied it in low-dimensional example systems. The strategy harnesses estimated local MFPTs from different initial configurations to a single target state. In the present work, we apply the optimal parameterization strategy to more challenging, high-dimensional molecular models, namely, synthetic molecular dynamics (MD) models of Trp-cage folding and unfolding, as well as atomistic MD models of NTL9 folding in high-friction and low-friction continuum solvents. In each system we use WE to estimate the MFPT for folding or unfolding events. We show that the optimal parameterization reduces the variance of MFPT estimates in three of four systems, with dramatic improvement in the most challenging atomistic system. Overall, the parameterization strategy improves the accuracy and reliability of WE estimates for the kinetics of biophysical processes.

physics.chem-ph

Democratizing Uncertainty Quantification

Uncertainty Quantification (UQ) is vital to safety-critical model-based analyses, but the widespread adoption of sophisticated UQ methods is limited by technical complexity. In this paper, we introduce UM-Bridge (the UQ and Modeling Bridge), a high-level abstraction and software protocol that facilitates universal interoperability of UQ software with simulation codes. It breaks down the technical complexity of advanced UQ applications and enables separation of concerns between experts. UM-Bridge democratizes UQ by allowing effective interdisciplinary collaboration, accelerating the development of advanced UQ methods, and making it easy to perform UQ analyses from prototype to High Performance Computing (HPC) scale. In addition, we present a library of ready-to-run UQ benchmark problems, all easily accessible through UM-Bridge. These benchmarks support UQ methodology research, enabling reproducible performance comparisons. We demonstrate UM-Bridge with several scientific applications, harnessing HPC resources even using UQ codes not designed with HPC support.

cs.MS

Featurizing Koopman Mode Decomposition For Robust Forecasting

This article introduces an advanced Koopman mode decomposition (KMD) technique -- coined Featurized Koopman Mode Decomposition (FKMD) -- that uses delay embedding and a learned Mahalanobis distance to enhance analysis and prediction of high dimensional dynamical systems. The delay embedding expands the observation space to better capture underlying manifold structure, while the Mahalanobis distance adjusts observations based on the system's dynamics. This aids in featurizing KMD in cases where good features are not a priori known. We show that FKMD improves predictions for a high-dimensional linear oscillator, a high-dimensional Lorenz attractor that is partially observed, and a cell signaling problem from cancer research.

math.DS

RiteWeight: Randomized Iterative Trajectory Reweighting for Steady-State Distributions Without Discretization Error

A significant challenge in molecular dynamics (MD) simulations is ensuring that sampled configurations converge to the equilibrium or nonequilibrium stationary distribution of interest. Lack of convergence constrains the estimation of free energies, rates, and mechanisms of complex molecular events. Here, we introduce the "Randomized ITErative trajectory reWeighting" (RiteWeight) algorithm to estimate a stationary distribution from unconverged simulation data. This method iteratively reweights trajectory segments in a self-consistent way by solving for the stationary distribution of a Markov state model (MSM), updating segment weights, and employing a new random clustering in each iteration. The iterative random clustering mitigates the phase-space discretization error inherent in existing trajectory reweighting techniques and yields quasi-continuous configuration-space distributions. We present mathematical analysis of the algorithm's fixed points as well as empirical validation using both synthetic MD Trp-cage trajectories, for which the stationary solution is exactly calculable, and standard atomistic MD Trp-cage trajectories extracted from a long reference simulation. In both test systems, we find that RiteWeight corrects flawed distributions and generates accurate observables for equilibrium and nonequilibrium steady states. The results highlight the value of correcting the underlying trajectory distribution rather than using a standard MSM

physics.comp-ph

Arbitrarily accurate, nonparametric coarse graining with Markov renewal processes and the Mori-Zwanzig formulation

Stochastic dynamics, such as molecular dynamics, are important in many scientific applications. However, summarizing and analyzing the results of such simulations is often challenging, due to the high dimension in which simulations are carried out, and consequently to the very large amount of data that is typically generated. Coarse graining is a popular technique for addressing this problem by providing compact and expressive representations. Coarse graining, however, potentially comes at the cost of accuracy, as dynamical information is in general lost when projecting the problem in a lower dimensional space. This article shows how to eliminate coarse-graining error using two key ideas. First, we represent coarse-grained dynamics as a Markov renewal process. Second, we outline a data-driven, non-parametric Mori-Zwanzig approach for computing jump times of the renewal process. Numerical tests on a small protein illustrate the method.

math.DS

A benchmark for the Bayesian inversion of coefficients in partial differential equations

Bayesian methods have been widely used in the last two decades to infer statistical properties of spatially variable coefficients in partial differential equations from measurements of the solutions of these equations. Yet, in many cases the number of variables used to parameterize these coefficients is large, and obtaining meaningful statistics of their values is difficult using simple sampling methods such as the basic Metropolis-Hastings (MH) algorithm -- in particular if the inverse problem is ill-conditioned or ill-posed. As a consequence, many advanced sampling methods have been described in the literature that converge faster than MH, for example by exploiting hierarchies of statistical models or hierarchies of discretizations of the underlying differential equation. At the same time, it remains difficult for the reader of the literature to quantify the advantages of these algorithms because there is no commonly used benchmark. This paper presents a benchmark Bayesian inverse problem -- namely, the determination of a spatially-variable coefficient, discretized by 64 values, in a Poisson equation, based on point measurements of the solution -- that fills the gap between widely used simple test cases (such as superpositions of Gaussians) and real applications that are difficult to replicate for developers of sampling algorithms. We provide a complete description of the test case, and provide an open source implementation that can serve as the basis for further experiments. We have also computed $2\times 10^{11}$ samples, at a cost of some 30 CPU years, of the posterior probability distribution from which we have generated detailed and accurate statistics against which other sampling algorithms can be tested.

math.NA

Unbiased estimation of equilibrium, rates, and committors from Markov state model analysis

Markov state models (MSMs) have been broadly adopted for analyzing molecular dynamics trajectories, but the approximate nature of the models that results from coarse-graining into discrete states is a long-known limitation. We show theoretically that, despite the coarse graining, in principle MSM-like analysis can yield unbiased estimation of key observables. We describe unbiased estimators for equilibrium state populations, for the mean first-passage time (MFPT) of an arbitrary process, and for state committors - i.e., splitting probabilities. Generically, the estimators are only asymptotically unbiased but we describe how extension of a recently proposed reweighting scheme can accelerate relaxation to unbiased values. Exactly accounting for 'sliding window' averaging over finite-length trajectories is a key, novel element of our analysis. In general, our analysis indicates that coarse-grained MSMs are asymptotically unbiased for steady-state properties only when appropriate boundary conditions (e.g., source-sink for MFPT estimation) are applied directly to trajectories, prior to calculation of the appropriate transition matrix.

physics.comp-ph

An ergodic theorem for the weighted ensemble method

We study weighted ensemble, an interacting particle method for sampling distributions of Markov chains that has been used in computational chemistry since the 1990s. Many important applications of weighted ensemble require the computation of long time averages. We establish the consistency of weighted ensemble in this setting by proving an ergodic theorem for time averages. As part of the proof, we derive explicit variance formulas that could be useful for optimizing the method.

math.NA

A splitting method to reduce MCMC variance

We explore whether splitting and killing methods can improve the accuracy of Markov chain Monte Carlo (MCMC) estimates of rare event probabilities, and we make three contributions. First, we prove that "weighted ensemble" is the only splitting and killing method that provides asymptotically consistent estimates when combined with MCMC. Second, we prove a lower bound on the asymptotic variance of weighted ensemble's estimates. Third, we give a constructive proof and numerical examples to show that weighted ensemble can approach this optimal variance bound, in many cases reducing the variance of MCMC estimates by multiple orders of magnitude.

math.NA

Optimizing weighted ensemble sampling of steady states

We propose parameter optimization techniques for weighted ensemble sampling of Markov chains in the steady-state regime. Weighted ensemble consists of replicas of a Markov chain, each carrying a weight, that are periodically resampled according to their weights inside of each of a number of bins that partition state space. We derive, from first principles, strategies for optimizing the choices of weighted ensemble parameters, in particular the choice of bins and the number of replicas to maintain in each bin. In a simple numerical example, we compare our new strategies with more traditional ones and with direct Monte Carlo.

math.NA

Transient probability currents provide upper and lower bounds on non-equilibrium steady-state currents in the Smoluchowski picture

Probability currents are fundamental in characterizing the kinetics of non-equilibrium processes. Notably, the steady-state current $J_{ss}$ for a source-sink system can provide the exact mean-first-passage time (MFPT) for the transition from source to sink. Because transient non-equilibrium behavior is quantified in some modern path sampling approaches, such as the "weighted ensemble" strategy, there is strong motivation to determine bounds on $J_{ss}$ -- and hence on the MFPT -- as the system evolves in time. Here we show that $J_{ss}$ is bounded from above and below by the maximum and minimum, respectively, of the current as a function of the spatial coordinate at any time $t$ for one-dimensional systems undergoing over-damped Langevin (i.e., Smoluchowski) dynamics and for higher-dimensional Smoluchowski systems satisfying certain assumptions when projected onto a single dimension. These bounds become tighter with time, making them of potential practical utility in a scheme for estimating $J_{ss}$ and the long-timescale kinetics of complex systems. Conceptually, the bounds result from the fact that extrema of the transient currents relax toward the steady-state current.

cond-mat.stat-mech

Generalizing Parallel Replica Dynamics: Trajectory Fragments, Asynchronous Computing, and PDMPs

We study the Parallel Replica Dynamics in a general setting. We introduce a trajectory fragment framework that can be used to design and prove consistency of Parallel Replica algorithms for generic Markov processes. We use our framework to formulate a novel condition that guarantees an asynchronous algorithm is consistent. Exploiting this condition and our trajectory fragment framework, we present new synchronous and asynchronous Parallel Replica algorithms for piecewise deterministic Markov processes.

math.NA

On the phase transition curve in a directed exponential random graph model

We consider a family of directed exponential random graph models parametrized by edges and outward stars. Much of the important statistical content of such models is given by the normalization constant of the models, and in particular, an appropriately scaled limit of the normalization, which is called the free energy. We derive precise asymptotics for the normalization constant for finite graphs. We use this to derive a formula for the free energy. The limit is analytic everywhere except along a curve corresponding to a first order phase transition. We examine unusual behavior of the model along the phase transition curve.

math.PR

Lattice Gas Models with Long Range Interactions

We study microcanonical lattice gas models with long range interactions, including power law interactions. We rigorously obtain a variational principle for the entropy. In a one dimensional example, we find a first order phase transition by proving the entropy is non-differentiable along a certain curve.

math-ph

Stationary averaging for multi-scale continuous time Markov chains using parallel replica dynamics

We propose two algorithms for simulating continuous time Markov chains in the presence of metastability. We show that the algorithms correctly estimate, under the ergodicity assumption, stationary averages of the process. Both algorithms, based on the idea of the parallel replica method, use parallel computing in order to explore metastable sets more efficiently. The algorithms require no assumptions on the Markov chains beyond ergodicity and the presence of identifiable metastability. In particular, there is no assumption on reversibility. For simpler illustration of the algorithms, we assume that a synchronous architecture is used throughout of the paper. We present error analyses, as well as numerical simulations on multi-scale stochastic reaction network models in order to demonstrate consistency of the method and its efficiency.

math.NA

Analysis and optimization of weighted ensemble sampling

We give a mathematical framework for weighted ensemble (WE) sampling, a binning and resampling technique for efficiently computing probabilities in molecular dynamics. We prove that WE sampling is unbiased in a very general setting that includes adaptive binning. We show that when WE is used for stationary calculations in tandem with a coarse model, the coarse model can be used to optimize the allocation of replicas in the bins.

math.NA

Asymptotic structure and singularities in constrained directed graphs

We study the asymptotics of large directed graphs, constrained to have certain densities of edges and/or outward $p$-stars. Our models are close cousins of exponential random graph models (ERGMs), in which edges and certain other subgraph densities are controlled by parameters. The idea of directly constraining edge and other subgraph densities comes from Radin and Sadun. Such modeling circumvents a phenomenon first made precise by Chatterjee and Diaconis: that in ERGMs it is often impossible to independently constrain edge and other subgraph densities. In all our models, we find that large graphs have either uniform or bipodal structure. When edge density (resp. $p$-star density) is fixed and $p$-star density (resp. edge density) is controlled by a parameter, we find phase transitions corresponding to a change from uniform to bipodal structure. When both edge and $p$-star density are fixed, we find only bipodal structures and no phase transition.

math.PR