Searcharxiv⌕ Search

arXiv subjects

Vincent A. Voelz

Publications and source records attributed to Vincent A. Voelz.

9 recordsLinked to original sources

Deep Generative Markov State Models with Experimental Restraints

A central challenge in molecular modeling is reconciling simulations with experimental observables, as force-field inaccuracies can distort equilibrium populations and long-timescale kinetics. While time-dependent restraints can improve simulated kinetics, time-dependent structural observables remain limited. Thus, most experimental data are time-averaged, informing equilibrium ensembles but not directly dynamics. Although ensemble refinement can improve thermodynamic accuracy, incorporating time-averaged observables into kinetic models remains difficult. Here, we introduce BICePs-reweighted Reversible DeepMSMs, combining Bayesian Inference of Conformational Populations (BICePs), variational learning of Markov processes, and maximum entropy (MaxEnt)/maximum caliber (MaxCal) principles to infer consistent thermodynamics and minimally perturbed kinetics. At its core is a reversible DeepMSM prior, obtained by applying an orthogonal transformation post-hoc to a pre-trained dynamics model (e.g., VAMPnet). This preserves the eigenspectrum exactly while yielding a valid transition matrix that is nonnegative, row-stochastic, and reversible. The reweighted model refines stationary populations, transition dynamics, and state-conditioned configurational landing densities. We validate the approach on a quadruple-well toy system and alanine dipeptide using backbone dihedral angles and J-coupling constants. In both systems, perturbing the observables induces predictable changes in the equilibrium ensemble and corresponding inferred kinetics. The resulting relaxation timescales and dynamical modes closely agree with analytical and MaxCal reference models. Furthermore, a diffusion-based generative model trained on MaxEnt landing densities produces physically realistic alanine dipeptide trajectories that reproduce thermodynamics and kinetics while satisfying the imposed experimental restraints.

physics.bio-ph↗

Automated optimization of force field parameters against ensemble-averaged measurements with Bayesian Inference of Conformational Populations

Accurate force fields are essential for reliable molecular simulations. These models are refined against quantum mechanical calculations and experimental measurements, which are subject to random and systematic errors. Bayesian Inference of Conformational Populations (BICePs) is a reweighting algorithm that reconciles simulated ensembles with sparse or noisy observables by sampling the full posterior distribution of conformational populations and experimental uncertainty. In this method, a metric called the BICePs score is used to perform model selection, by calculating the free energy of "turning on" the conformational populations under experimental restraints. This approach, when used with improved likelihood functions to deal with experimental outliers, can be used for force field validation (Raddi et al. 2025). Here, we extend the BICePs approach to perform automated force field refinement while simultaneously sampling the full distribution of uncertainties, using a variational method to minimize the BICePs score. To demonstrate the utility of this method, we refine multiple interaction parameters for a 12-mer HP lattice model using ensemble-averaged distance measurements as restraints. To illustrate the resilience of BICePs in the presence of unknown random and systematic errors, we assess the performance of our algorithm through repeated optimizations and under various extents of experimental error. Our results suggest that variational optimization of the BICePs score is a promising direction for robust and automatic parameterization of molecular potentials.

physics.chem-ph↗

Automatic Forward Model Parameterization with Bayesian Inference of Conformational Populations

To quantify how well theoretical predictions of structural ensembles agree with experimental measurements, we depend on the accuracy of forward models. These models are computational frameworks that generate observable quantities from molecular configurations based on empirical relationships linking specific molecular properties to experimental measurements. Bayesian Inference of Conformational Populations (BICePs) is a reweighting algorithm that reconciles simulated ensembles with ensemble-averaged experimental observations, even when such observations are sparse and/or noisy. This is achieved by sampling the posterior distribution of conformational populations under experimental restraints as well as sampling the posterior distribution of uncertainties due to random and systematic error. In this study, we enhance the algorithm for the refinement of empirical forward model (FM) parameters. We introduce and evaluate two novel methods for optimizing FM parameters. The first method treats FM parameters as nuisance parameters, integrating over them in the full posterior distribution. The second method employs variational minimization of a quantity called the BICePs score that reports the free energy of `turning on` the experimental restraints. This technique, coupled with improved likelihood functions for handling experimental outliers, facilitates force field validation and optimization, as illustrated in recent studies (Raddi et al. 2023, 2024). Using this approach, we refine parameters that modulate the Karplus relation, crucial for accurate predictions of J-coupling constants based on dihedral angles between interacting nuclei. We validate this approach first with a toy model system, and then for human ubiquitin, predicting six sets of Karplus parameters. Finally, we demonstrate that our framework naturally generalizes optimization to any differentiable forward model...

physics.bio-ph↗

Structure-Based Experimental Datasets for Benchmarking Protein Simulation Force Fields

This review article provides an overview of structurally oriented experimental datasets that can be used to benchmark protein force fields, focusing on data generated by nuclear magnetic resonance (NMR) spectroscopy and room temperature (RT) protein crystallography. We discuss what the observables are, what they tell us about structure and dynamics, what makes them useful for assessing force field accuracy, and how they can be connected to molecular dynamics simulations carried out using the force field one wishes to benchmark. We also touch on statistical issues that arise when comparing simulations with experiment. We hope this article will be particularly useful to computational researchers and trainees who develop, benchmark, or use protein force fields for molecular simulations.

q-bio.BM↗

Folding@home: achievements from over twenty years of citizen science herald the exascale era

Simulations of biomolecules have enormous potential to inform our understanding of biology but require extremely demanding calculations. For over twenty years, the Folding@home distributed computing project has pioneered a massively parallel approach to biomolecular simulation, harnessing the resources of citizen scientists across the globe. Here, we summarize the scientific and technical advances this perspective has enabled. As the project's name implies, the early years of Folding@home focused on driving advances in our understanding of protein folding by developing statistical methods for capturing long-timescale processes and facilitating insight into complex dynamical processes. Success laid a foundation for broadening the scope of Folding@home to address other functionally relevant conformational changes, such as receptor signaling, enzyme dynamics, and ligand binding. Continued algorithmic advances, hardware developments such as GPU-based computing, and the growing scale of Folding@home have enabled the project to focus on new areas where massively parallel sampling can be impactful. While previous work sought to expand toward larger proteins with slower conformational changes, new work focuses on large-scale comparative studies of different protein sequences and chemical compounds to better understand biology and inform the development of small molecule drugs. Progress on these fronts enabled the community to pivot quickly in response to the COVID-19 pandemic, expanding to become the world's first exascale computer and deploying this massive resource to provide insight into the inner workings of the SARS-CoV-2 virus and aid the development of new antivirals. This success provides a glimpse of what's to come as exascale supercomputers come online, and Folding@home continues its work.

q-bio.BM↗

Assigning Confidence to Molecular Property Prediction

Introduction: Computational modeling has rapidly advanced over the last decades, especially to predict molecular properties for chemistry, material science and drug design. Recently, machine learning techniques have emerged as a powerful and cost-effective strategy to learn from existing datasets and perform predictions on unseen molecules. Accordingly, the explosive rise of data-driven techniques raises an important question: What confidence can be assigned to molecular property predictions and what techniques can be used for that purpose? Areas covered: In this work, we discuss popular strategies for predicting molecular properties relevant to drug design, their corresponding uncertainty sources and methods to quantify uncertainty and confidence. First, our considerations for assessing confidence begin with dataset bias and size, data-driven property prediction and feature design. Next, we discuss property simulation via molecular docking, and free-energy simulations of binding affinity in detail. Lastly, we investigate how these uncertainties propagate to generative models, as they are usually coupled with property predictors. Expert opinion: Computational techniques are paramount to reduce the prohibitive cost and timing of brute-force experimentation when exploring the enormous chemical space. We believe that assessing uncertainty in property prediction models is essential whenever closed-loop drug design campaigns relying on high-throughput virtual screening are deployed. Accordingly, considering sources of uncertainty leads to better-informed experimental validations, more reliable predictions and to more realistic expectations of the entire workflow. Overall, this increases confidence in the predictions and designs and, ultimately, accelerates drug design.

cs.LG↗

Adaptive Markov State Model estimation using short reseeding trajectories

In the last decade, advances in molecular dynamics (MD) and Markov State Model (MSM) methodologies have made possible accurate and efficient estimation of kinetic rates and reactive pathways for complex biomolecular dynamics occurring on slow timescales. A promising approach to enhanced sampling of MSMs is to use so-called "adaptive" methods, in which new MD trajectories are "seeded" preferentially from previously identified states. Here, we investigate the performance of various MSM estimators applied to reseeding trajectory data, for both a simple 1D free energy landscape, and for mini-protein folding MSMs of WW domain and NTL9(1-39). Our results reveal the practical challenges of reseeding simulations, and suggest a simple way to reweight seeding trajectory data to better estimate both thermodynamic and kinetic quantities.

q-bio.BM↗

Is the free energy landscape informative about transition rates? Lessons from the kinetic Ising model

An oft-used concept in modeling macromolecules is the free energy landscape, obtained by coarse-graining a vast number of microstates into a low-dimensional mesh of mesostates. If the landscape contains two or more local minima (macrostates),one can compute global rate constants provided the dynamics of the dividing barrier regions are known. Here we compared experimental rate constants between ordered states in a kinetic Ising model with rates calculated from a coarse-grained master equation derived from the microcanonical ensemble. The coarse-grained macroscopic rate constants were roughly 50 % larger than experiment across a range of environmental constraints, suggesting a systematic impediment of configurational progress on the microscopic scale that is specific to the structure of the Ising model. The error in coarse-graining lay with the calculation of the diffusion coefficient rather than with the shape of the free energy landscape, as ensemble- and time-averaged estimates of the latter were indistinguishable. Fluctuation analysis in the form of Nyquist theorem also failed to substantially improve the value of the effective diffusion coefficient, suggesting a failure of the fluctuation-dissipation theorem. These findings from the Ising model raises doubts over the validity of the free energy landscape approach in calculating absolute transition rates for more complex systems such as proteins.

cond-mat.stat-mech↗

A maximum-caliber approach to predicting perturbed folding kinetics due to mutations

We present a maximum-caliber method for inferring transition rates of a Markov State Model (MSM) with perturbed equilibrium populations, given estimates of state populations and rates for an unperturbed MSM. It is similar in spirit to previous approaches but given the inclusion of prior information it is more robust and simple to implement. We examine its performance in simple biased diffusion models of kinetics, and then apply the method to predicting changes in folding rates for several highly non-trivial protein folding systems for which non-native interactions play a significant role, including (1) tryptophan variants of GB1 hairpin, (2) salt-bridge mutations of Fs peptide helix, and (3) MSMs built from ultra-long folding trajectories of FiP35 and GTT variants of WW domain. In all cases, the method correctly predicts changes in folding rates, suggesting the wide applicability of maximum-caliber approaches to efficiently predict how mutations perturb protein conformational dynamics.

q-bio.BM↗