SearcharxivSearch

arXiv subjects

Jakub Rydzewski

Publications and source records attributed to Jakub Rydzewski.

At least 19 recordsLinked to original sources

Unraveling the Mechanism of Drug Binding to SARS-CoV-2 RNA Pseudoknot with Thermodynamics-Driven Machine Learning

The pseudoknot secondary structure in SARS-CoV-2 RNA is essential for regulating protein synthesis through $-$1 programmed ribosomal frameshifting ($-1$ PRF), a mechanism that allows the virus to generate both structural and non-structural proteins from overlapping reading frames. This pseudoknot exhibits both threaded and unthreaded long-lived topologies. The influence of ligand binding on its folding is a process critical for the development of $-$1 PRF small-molecule inhibitors. Understanding this process through unbiased molecular dynamics (MD) simulations can be facilitated by introducing collective variables (CVs) that capture the corresponding slowest dynamical modes. Here, we use spectral map (SM), a thermodynamics-driven machine learning technique, to learn such CVs directly from all-atom MD trajectories of the SARS-CoV-2 RNA pseudoknot in complex with the $-$1 PRF inhibitor merafloxacin and its two structural analogs in neutral and ionized forms. Free-energy landscapes (FELs) derived from the learned CVs indicate that ligand-induced destabilization is topology-selective. In the threaded pseudoknot, the inhibitors destabilize the S2 stem, while in the unthreaded pseudoknot, destabilization occurs in the S1 and S3 stems. Furthermore, the extent to which each ligand reshapes the FEL matches experimentally reported antiviral potency, whereas the protonation state qualitatively alters dynamics within the same RNA topology. Overall, our results show how pseudoknot topology, ligand type, and protonation state collectively influence the slow conformational dynamics of viral RNA and establish physiological protonation as a critical factor for modeling RNA-targeted drug action.

physics.bio-ph

Constructing Generalized Sample Transition Probabilities with Biased Simulations

In molecular dynamics (MD) simulations, accessing transition probabilities between states is crucial for understanding kinetic information, such as reaction paths and rates. However, standard MD simulations are hindered by the capacity to visit the states of interest, prompting the use of enhanced sampling to accelerate the process. Unfortunately, biased simulations alter the inherent probability distributions, making kinetic computations using techniques such as diffusion maps challenging. Here, we use a coarse-grained Markov chain to estimate the intrinsic pairwise transition probabilities between states sampled from a biased distribution. Our method, which we call the generalized sample transition probability (GSTP), can recover transition probabilities without relying on an underlying stochastic process and specifying the form of the kernel function, which is necessary for the diffusion map method. The proposed algorithm is validated on model systems such as a harmonic oscillator, alanine dipeptide in vacuum, and met-enkephalin in solvent. The results demonstrate that GSTP effectively recovers the unbiased eigenvalues and eigenstates from biased data. GSTP provides a general framework for analyzing kinetic information in complex systems, where biased simulations are necessary to access longer timescales.

physics.chem-ph

NeuralTSNE: A Python Package for the Dimensionality Reduction of Molecular Dynamics Data Using Neural Networks

Unsupervised machine learning has recently gained much attention in the field of molecular dynamics (MD). Particularly, dimensionality reduction techniques have been regularly employed to analyze large volumes of high-dimensional MD data to gain insight into hidden information encoded in MD trajectories. Among many such techniques, t-distributed stochastic neighbor embedding (t-SNE) is particularly popular. A parametric version of t-SNE that employs neural networks is less commonly known, yet it has demonstrated superior performance in dimensionality reduction compared to the standard implementation. Here, we present a Python package called NeuralTSNE with our implementation of parametric t-SNE. The implementation is done using the PyTorch library and the PyTorch Lightning framework and can be imported as a module or used from the command line. We show that NeuralTSNE offers an easy-to-use tool for the analysis of MD data.

physics.chem-ph

Illuminating Protein Dynamics: A Review of Computational Methods for Studying Photoactive Proteins

Photoactive proteins absorb light and undergo structural changes that enable them to perform essential biological functions. These proteins are critical for understanding light-induced biological processes, making them important in biophysics, biotechnology, and medicine. One effective approach to uncovering photoactive processes is through computational methods. These techniques provide atomic-level insights into the structural, electronic, and dynamic changes that occur upon light absorption. By employing these methods, we can gain a better understanding of processes that are challenging to capture experimentally, such as chromophore isomerization and protein conformational changes. Here, we provide a brief overview of the different families of photoactive proteins and the computational methods used to study them, including bioinformatics, molecular dynamics, and enhanced sampling. Our review can serve as an introduction to computational methods for studying light-activated molecular processes, specifically targeting researchers beginning their journey in this field.

physics.chem-ph

Decoding Binding Pathways of Ligands in Prolyl Oligopeptidase

Neurodegenerative diseases, such as Alzheimer's and Parkinson's, pose a growing global health burden. Prolyl oligopeptidase (PREP) has emerged as a potential therapeutic target in these diseases. Recent studies have shown that direct interaction between PREP and pathological proteins, such as $\alpha$-synuclein and Tau, influences protein aggregation and neuronal function. While most known PREP inhibitors primarily target its enzymatic functions, a new class of ligands, known as HUPs, specifically modulates protein-protein interactions (PPIs), which are crucial in neurodegenerative diseases. These structurally distinct ligands exhibit diverse binding behaviors, highlighting the importance of understanding their binding pathways. In this study, we analyzed the binding pathways and stability of diverse ligands using molecular dynamics simulations and enhanced sampling techniques. Traditional inhibitors, such as KYP-2047, target the active site between the catalytic domains of PREP and the $\beta$-propeller domain, while HUP ligands bind to alternative regions, such as the hinge site, potentially disrupting non-enzymatic PPIs. We demonstrated that structural variations among ligands lead to distinct binding and unbinding pathways. Free-energy profiles from umbrella sampling revealed key kinetic bottlenecks and differences in pathways. For example, HUP-55 exhibits pathway hopping, characterized by diffuse exploration of binding regions before selecting an exit, while KYP-2047 prefers the central tunnel of the $\beta$-propeller domain even under perturbations. These results suggest that the dynamic interaction between ligands and PREP plays a critical role in their mechanism. The ability of HUPs to interact with multiple binding sites and adapt to PREP's conformational changes may be essential for their PPI-targeting effects.

physics.chem-ph

Machine Learning of Slow Collective Variables and Enhanced Sampling via Spatial Techniques

Understanding the long-time dynamics of complex physical processes depends on our ability to recognize patterns. To simplify the description of these processes, we often introduce a set of reaction coordinates, customarily referred to as collective variables (CVs). The quality of these CVs heavily impacts our comprehension of the dynamics, often influencing the estimates of thermodynamics and kinetics from atomistic simulations. Consequently, identifying CVs poses a fundamental challenge in chemical physics. Recently, significant progress was made by leveraging the predictive ability of unsupervised machine learning techniques to determine CVs. Many of these techniques require temporal information to learn slow CVs that correspond to the long timescale behavior of the studied process. Here, however, we specifically focus on techniques that can identify CVs corresponding to the slowest transitions between states without needing temporal trajectories as input, instead using the spatial characteristics of the data. We discuss the latest developments in this category of techniques and briefly discuss potential directions for thermodynamics-informed spatial learning of slow CVs.

physics.chem-ph

A Note on Spectral Map

In molecular dynamics (MD) simulations, transitions between states are often rare events due to energy barriers that exceed the thermal temperature. Because of their infrequent occurrence and the huge number of degrees of freedom in molecular systems, understanding the physical properties that drive rare events is immensely difficult. A common approach to this problem is to propose a collective variable (CV) that describes this process by a simplified representation. However, choosing CVs is not easy, as it often relies on physical intuition. Machine learning (ML) techniques provide a promising approach for effectively extracting optimal CVs from MD data. Here, we provide a note on a recent unsupervised ML method called spectral map, which constructs CVs by maximizing the timescale separation between slow and fast variables in the system.

physics.chem-ph

PLUMED Tutorials: a collaborative, community-driven learning ecosystem

In computational physics, chemistry, and biology, the implementation of new techniques in a shared and open source software lowers barriers to entry and promotes rapid scientific progress. However, effectively training new software users presents several challenges. Common methods like direct knowledge transfer and in-person workshops are limited in reach and comprehensiveness. Furthermore, while the COVID-19 pandemic highlighted the benefits of online training, traditional online tutorials can quickly become outdated and may not cover all the software's functionalities. To address these issues, here we introduce ``PLUMED Tutorials'', a collaborative model for developing, sharing, and updating online tutorials. This initiative utilizes repository management and continuous integration to ensure compatibility with software updates. Moreover, the tutorials are interconnected to form a structured learning path and are enriched with automatic annotations to provide broader context. This paper illustrates the development, features, and advantages of PLUMED Tutorials, aiming to foster an open community for creating and sharing educational resources.

physics.ed-ph

Spectral Map for Slow Collective Variables, Markovian Dynamics, and Transition State Ensembles

Understanding the behavior of complex molecular systems is a fundamental problem in physical chemistry. To describe the long-time dynamics of such systems, which is responsible for their most informative characteristics, we can identify a few slow collective variables (CVs) while treating the remaining fast variables as thermal noise. This enables us to simplify the dynamics and treat it as diffusion in a free-energy landscape spanned by slow CVs, effectively rendering the dynamics Markovian. Our recent statistical learning technique, spectral map [Rydzewski, J. Phys. Chem. Lett. 2023, 14, 22, 5216-5220], explores this strategy to learn slow CVs by maximizing a spectral gap of a transition matrix. In this work, we introduce several advancements into our framework, using a high-dimensional reversible folding process of a protein as an example. We implement an algorithm for coarse-graining Markov transition matrices to partition the reduced space of slow CVs kinetically and use it to define a transition state ensemble. We show that slow CVs learned by spectral map closely approach the Markovian limit for an overdamped diffusion. We demonstrate that coordinate-dependent diffusion coefficients only slightly affect the constructed free-energy landscapes. Finally, we present how spectral map can be used to quantify the importance of features and compare slow CVs with structural descriptors commonly used in protein folding. Overall, we demonstrate that a single slow CV learned by spectral map can be used as a physical reaction coordinate to capture essential characteristics of protein folding.

physics.chem-ph

Spectral Maps for Learning Reduced Representations of Molecular Systems

Investigating processes in complex molecular systems, which are characterized by many variables, is a crucial problem in computational physics. These systems can be reduced to a few meaningful degrees of freedom known as collective variables (CVs). However, identifying these CVs is a significant challenge, especially for systems with long-lived metastable states. This is because the information about the slow kinetics of rare transitions needs to be encoded in CVs. In this talk, we review recent advances in learning slow CVs and focus mainly on our spectral map technique, a promising deep-learning method that learns CVs based on the slowest timescales. By maximizing the spectral gap between slow and fast eigenvalues of a Markov transition matrix constructed from simulation data, our method effectively captures a simplified representation of alanine dipeptide in solvent. This practical application of our method demonstrates its ability to extract slow CVs, making it a valuable tool for analyzing complex systems.

physics.chem-ph

Reweighted Manifold Learning of Collective Variables from Enhanced Sampling Simulations

Enhanced sampling methods are indispensable in computational physics and chemistry, where atomistic simulations cannot exhaustively sample the high-dimensional configuration space of dynamical systems due to the sampling problem. A class of such enhanced sampling methods works by identifying a few slow degrees of freedom, termed collective variables (CVs), and enhancing the sampling along these CVs. Selecting CVs to analyze and drive the sampling is not trivial and often relies on physical and chemical intuition. Despite routinely circumventing this issue using manifold learning to estimate CVs directly from standard simulations, such methods cannot provide mappings to a low-dimensional manifold from enhanced sampling simulations as the geometry and density of the learned manifold are biased. Here, we address this crucial issue and provide a general reweighting framework based on anisotropic diffusion maps for manifold learning that takes into account that the learning data set is sampled from a biased probability distribution. We consider manifold learning methods based on constructing a Markov chain describing transition probabilities between high-dimensional samples. We show that our framework reverts the biasing effect yielding CVs that correctly describe the equilibrium density. This advancement enables the construction of low-dimensional CVs using manifold learning directly from data generated by enhanced sampling simulations. We call our framework reweighted manifold learning. We show that it can be used in many manifold learning techniques on data from both standard and enhanced sampling simulations.

physics.chem-ph

Selecting High-Dimensional Representations of Physical Systems by Reweighted Diffusion Maps

Constructing reduced representations of high-dimensional systems is a fundamental problem in physical chemistry. Many unsupervised machine learning methods can automatically find such low-dimensional representations. However, an often overlooked problem is what high-dimensional representation should be used to describe systems before dimensionality reduction. Here, we address this issue using a recently developed method called reweighted diffusion map [J. Chem. Theory Comput. 2022, 18, 7179-7192]. We show how high-dimensional representations can be quantitatively selected by exploring the spectral decomposition of Markov transition matrices built from data obtained from standard or enhanced sampling atomistic simulations. We demonstrate the performance of the method in several high-dimensional examples.

physics.chem-ph

Spectral Map: Embedding Slow Kinetics in Collective Variables

The dynamics of physical systems that require high-dimensional representation can often be captured in a few meaningful degrees of freedom called collective variables (CVs). However, identifying CVs is challenging and constitutes a fundamental problem in physical chemistry. This problem is even more pronounced when CVs information about slow kinetics related to rare transitions between long-lived metastable states. To address this issue, we propose an unsupervised deep-learning method called spectral map. Our method constructs slow CVs by maximizing the spectral gap between slow and fast eigenvalues of a transition matrix estimated by an anisotropic diffusion kernel. We demonstrate our method in several high-dimensional reversible folding processes.

physics.chem-ph

Learning Markovian Dynamics with Spectral Maps

The long-time behavior of many complex molecular systems is often governed by slow relaxation dynamics that can be described by a few reaction coordinates referred to as collective variables (CVs). However, identifying CVs hidden in a high-dimensional configuration space poses a fundamental challenge in chemical physics. To address this problem, we expand on a recently introduced deep-learning technique called spectral map [Rydzewski, J. Phys. Chem. Lett. 2023, 14, 22, 5216-5220]. Spectral map learns CVs by maximizing a spectral gap between slow and fast eigenvalues of a Markov transition matrix describing anisotropic diffusion. An introduced modification in the learning algorithm allows spectral map to represent multiscale free-energy landscapes. Through a Markov state model analysis, we validate that spectral map learns slow CVs related to the dominant relaxation timescales and discerns between long-lived metastable states.

physics.chem-ph

Manifold Learning in Atomistic Simulations: A Conceptual Review

Analyzing large volumes of high-dimensional data requires dimensionality reduction: finding meaningful low-dimensional structures hidden in their high-dimensional observations. Such practice is needed in atomistic simulations of complex systems where even thousands of degrees of freedom are sampled. An abundance of such data makes gaining insight into a specific physical problem strenuous. Our primary aim in this review is to focus on unsupervised machine learning methods that can be used on simulation data to find a low-dimensional manifold providing a collective and informative characterization of the studied process. Such manifolds can be used for sampling long-timescale processes and free-energy estimation. We describe methods that can work on datasets from standard and enhanced sampling atomistic simulations. Unlike recent reviews on manifold learning for atomistic simulations, we consider only methods that construct low-dimensional manifolds based on Markov transition probabilities between high-dimensional samples. We discuss these techniques from a conceptual point of view, including their underlying theoretical frameworks and possible limitations.

physics.comp-ph

Multiscale reweighted stochastic embedding (MRSE): Deep learning of collective variables for enhanced sampling

Machine learning methods provide a general framework for automatically finding and representing the essential characteristics of simulation data. This task is particularly crucial in enhanced sampling simulations. There we seek a few generalized degrees of freedom, referred to as collective variables (CVs), to represent and drive the sampling of the free energy landscape. In theory, these CVs should separate different metastable states and correspond to the slow degrees of freedom of the studied physical process. To this aim, we propose a new method that we call multiscale reweighted stochastic embedding (MRSE). Our work builds upon a parametric version of stochastic neighbor embedding. The technique automatically learns CVs that map a high-dimensional feature space to a low-dimensional latent space via a deep neural network. We introduce several new advancements to stochastic neighbor embedding methods that make MRSE especially suitable for enhanced sampling simulations: (1) weight-tempered random sampling as a landmark selection scheme to obtain training data sets that strike a balance between equilibrium representation and capturing important metastable states lying higher in free energy; (2) a multiscale representation of the high-dimensional feature space via a Gaussian mixture probability model; and (3) a reweighting procedure to account for training data from a biased probability distribution. We show that MRSE constructs low-dimensional CVs that can correctly characterize the different metastable states in three model systems: the Müller-Brown potential, alanine dipeptide, and alanine tetrapeptide.

physics.chem-ph

maze: Heterogeneous Ligand Unbinding along Transient Protein Tunnels

Recent developments in enhanced sampling methods showed that it is possible to reconstruct ligand unbinding pathways with spatial and temporal resolution inaccessible to experiments. Ideally, such techniques should provide an atomistic definition of possibly many reaction pathways, because crude estimates may lead either to overestimating energy barriers, or inability to sample hidden energy barriers that are not captured by reaction pathway estimates. Here we provide an implementation of a new method [J. Rydzewski \& O. Valsson, J. Chem. Phys. {\bf 150}, 221101 (2019)] dedicated entirely to sampling the reaction pathways of the ligand-protein dissociation process. The program, called \texttt{maze}, is implemented as an official module for PLUMED 2, an open source library for enhanced sampling in molecular systems, and comprises algorithms to find multiple heterogeneous reaction pathways of ligand unbinding from proteins during atomistic simulations. The \texttt{maze} module requires only a crystallographic structure to start a simulation, and does not depend on many \textit{ad hoc} parameters. The program is based on enhanced sampling and non-convex optimization methods. To present its applicability and flexibility, we provide several examples of ligand unbinding pathways along transient protein tunnels reconstructed by \texttt{maze} in a model ligand-protein system, and discuss the details of the implementation.

physics.bio-ph

Memetic Algorithms for Ligand Expulsion from Protein Cavities

Ligand diffusion through proteins is a fundamental process governing biological signaling and enzymatic catalysis. The complex topology of protein tunnels results in difficulties with computing ligand escape pathways by standard molecular dynamics (MD) simulations. Here, two novel methods for searching of ligand exit pathways and cavity exploration are proposed: memory random acceleration MD (mRAMD), and memetic algorithms (MA). In mRAMD, finding exit pathways is based on a non-Markovian biasing that is introduced to optimize the unbinding force. In MA, hybrid learning protocols are exploited to predict optimal ligand exit paths. The methods are tested on three proteins with increasing complexity of tunnels: M2 muscarinic receptor, nitrile hydratase, and cytochrome P450cam. In these cases, the proposed methods outperform standard techniques that are used currently to find ligand egress pathways. The proposed approach is general and appropriate for accelerated transport of an object through a network of protein tunnels.

physics.chem-ph