SearcharxivSearch

arXiv subjects

Daniel T. Zhang

Publications and source records attributed to Daniel T. Zhang.

8 recordsLinked to original sources

Generalized Path Reweighting and History-Dependent Free Energies

Transition interface sampling (TIS) and replica exchange TIS (RETIS) are powerful methods for computing rates of rare events inaccessible to straightforward molecular dynamics (MD) simulations. Path reweighting extends their output, enabling the evaluation of diverse thermodynamic and kinetic quantities, including reaction prediction metrics, activation barriers, committor functions, and free energies. The recently developed Infinity-RETIS algorithm boosts parallel efficiency through asynchronous replica exchanges in the infinite-swap limit, eliminating the wall-time bottlenecks of conventional RETIS. This approach introduces fractional samples and biased sampling distributions, requiring a generalized path reweighting framework, for which we derive expressions demonstrating how exact dynamic and thermodynamic variables can be computed. We then focus on a special class of free energy surfaces defined by history-dependent conditions, whose values are influenced by kinetic factors such as particle mass and friction, unlike standard unconditional free energy surfaces. Even with suboptimal reaction coordinates, these conditional free energies can reveal kinetically relevant barriers that may be misrepresented by standard unconditional free energies, thereby providing a rigorous and versatile tool for characterizing complex molecular transitions.

physics.chem-ph

Estimating Full Path Lengths and Kinetics from Partial Path Transition Interface Sampling Simulations

Assessing the time scale of biological processes using molecular dynamics (MD) simulations with sufficient statistical accuracy is a challenging task, as processes are often rare and/or slow events, which may extend largely beyond the time scale of what is accessible with modern day high performance computational infrastructure. Recently, the replica exchange partial path transition interface sampling (REPPTIS) algorithm was developed to study rare and slow events involving metastable states along their reactive pathways. REPPTIS is a path sampling method where paths are cut short to reduce the computational cost, while combining this with the efficiency offered by replica exchange between the partial path ensembles. However, REPPTIS still lacks a formalism to extract time-dependent properties, such as mean first passage times, fluxes, and rates, from the short partial paths. In this work, we introduce a Markov state model (MSM) framework to estimate full path lengths and kinetic properties from the overlapping partial paths generated by REPPTIS. The framework results in newly derived closed formulas for the REPPTIS crossing probability, mean first passage times (MFPTs), flux, and rate constant. Our approach is then validated using simulations of Brownian and Langevin particles on a series of one-dimensional potential energy profiles as well as the dissociation of KCl in solution, demonstrating that REPPTIS accurately reproduces the exact kinetics benchmark. The MSM framework is further applied to the trypsin-benzamidine complex to compute the dissociation rate as a test case of a biological system, albeit the computed rate underestimates the experimental value. In conclusion, our MSM framework equips REPPTIS simulations with a robust theoretical and practical foundation for extracting kinetic information from computationally efficient partial paths.

physics.comp-ph

Path sampling challenges in large biomolecular systems: RETIS and REPPTIS for ABL-imatinib kinetics

Predicting the kinetics of drug-protein interactions is crucial for understanding drug efficacy, particularly in personalized medicine, where protein mutations can significantly alter drug residence times. This study applies Replica Exchange Transition Interface Sampling (RETIS) and its Partial Path variant (REPPTIS) to investigate the dissociation kinetics of imatinib from Abelson nonreceptor tyrosine kinase (ABL) and mutants relevant to chronic myeloid leukemia therapy. These path-sampling methods offer a bias-free alternative to conventional approaches requiring qualitative predefined reaction coordinates. Nevertheless, the complex free-energy landscape of ABL-imatinib dissociation presents significant challenges. Multiple metastable states and orthogonal barriers lead to parallel unbinding pathways, complicating convergence in TIS-based methods. Despite employing computational efficiency strategies such as asynchronous replica exchange, full convergence remained elusive. This work provides a critical assessment of path sampling in high-dimensional biological systems, discussing the need for enhanced initialization strategies, advanced Monte Carlo path generation moves, and machine learning-derived reaction coordinates to improve kinetic predictions of drug dissociation with minimal prior knowledge.

physics.bio-ph

Enhanced path sampling using subtrajectory Monte Carlo moves

Path sampling allows the study of rare events like chemical reactions, nucleation and protein folding via a Monte Carlo (MC) exploration in path space. Instead of configuration points, this method samples short molecular dynamics (MD) trajectories with specific start- and end-conditions. As in configuration MC, its efficiency highly depends on the types of MC moves. Since the last two decades, the central MC move for path sampling has been the so-called shooting move in which a perturbed phase point of the old path is propagated backward and forward in time to generate a new path. Recently, we proposed the subtrajectory moves, stone-skipping (SS) and web-throwing (WT), that are demonstrably more efficient. However, the one-step crossing requirement makes them somewhat more difficult to implement in combination with external MD programs or when the order parameter determination is expensive. In this article, we present strategies to address the issue. The most generic solution is a new member of subtrajectory moves, wire fencing (WF), that is less thrifty than the SS, but more versatile. This makes it easier to link path sampling codes with external MD packages and provides a practical solution for cases where the calculation of the order parameter is expensive or not a simple function of geometry. We demonstrate the WF move in a double well Langevin model, a thin film breaking transition based on classical force fields, and a smaller ruthenium redox reaction at the ab initio level in which the order parameter explicitly depends on the electron density.

physics.chem-ph

Exchanging replicas with unequal cost, infinitely and permanently

We developed a replica exchange method that is effectively parallelizable even if the computational cost of the Monte Carlo moves in the parallel replicas are considerably different, for instance, because the replicas run on different type of processor units or because of the algorithmic complexity. To prove detailed-balance, we make a paradigm shift from the common conceptual viewpoint in which the set of parallel replicas represents a high-dimensional superstate, to an ensemble based criterion in which the other ensembles represent an environment that might or might not participate in the Monte Carlo move. In addition, based on a recent algorithm for computing permanents, we effectively increase the exchange rate to infinite without the steep factorial scaling as function of the number of replicas. We illustrate the effectiveness of the replica exchange methodology by combining it with a quantitative path sampling method, replica exchange transition interface sampling (RETIS), in which the costs for a Monte Carlo move can vary enormously as paths in a RETIS algorithm do not have the same length and the average path lengths tend to vary considerably for the different path ensembles that run in parallel. This combination, coined $\infty$RETIS, was tested on three model systems.

physics.comp-ph

Online Boosting for Multilabel Ranking with Top-k Feedback

We present online boosting algorithms for multilabel ranking with top-k feedback, where the learner only receives information about the top k items from the ranking it provides. We propose a novel surrogate loss function and unbiased estimator, allowing weak learners to update themselves with limited information. Using these techniques we adapt full information multilabel ranking algorithms (Jung and Tewari, 2018) to the top-k feedback setting and provide theoretical performance bounds which closely match the bounds of their full information counterparts, with the cost of increased sample complexity. These theoretical results are further substantiated by our experiments, which show a small gap in performance between the algorithms for the top-k feedback setting and that for the full information setting across various datasets.

stat.ML

Online Multiclass Boosting with Bandit Feedback

We present online boosting algorithms for multiclass classification with bandit feedback, where the learner only receives feedback about the correctness of its prediction. We propose an unbiased estimate of the loss using a randomized prediction, allowing the model to update its weak learners with limited information. Using the unbiased estimate, we extend two full information boosting algorithms (Jung et al., 2017) to the bandit setting. We prove that the asymptotic error bounds of the bandit algorithms exactly match their full information counterparts. The cost of restricted feedback is reflected in the larger sample complexity. Experimental results also support our theoretical findings, and performance of the proposed models is comparable to that of an existing bandit boosting algorithm, which is limited to use binary weak learners.

stat.ML

A Data Science Approach to Understanding Residential Water Contamination in Flint

When the residents of Flint learned that lead had contaminated their water system, the local government made water-testing kits available to them free of charge. The city government published the results of these tests, creating a valuable dataset that is key to understanding the causes and extent of the lead contamination event in Flint. This is the nation's largest dataset on lead in a municipal water system. In this paper, we predict the lead contamination for each household's water supply, and we study several related aspects of Flint's water troubles, many of which generalize well beyond this one city. For example, we show that elevated lead risks can be (weakly) predicted from observable home attributes. Then we explore the factors associated with elevated lead. These risk assessments were developed in part via a crowd sourced prediction challenge at the University of Michigan. To inform Flint residents of these assessments, they have been incorporated into a web and mobile application funded by \texttt{Google.org}. We also explore questions of self-selection in the residential testing program, examining which factors are linked to when and how frequently residents voluntarily sample their water.

cs.LG