SearcharxivSearch

arXiv subjects

Debsindhu Bhowmik

Publications and source records attributed to Debsindhu Bhowmik.

8 recordsLinked to original sources

Transferring a molecular foundation model for polymer property predictions

Transformer-based large language models have remarkable potential to accelerate design optimization for applications such as drug development and materials discovery. Self-supervised pretraining of transformer models requires large-scale datasets, which are often sparsely populated in topical areas such as polymer science. State-of-the-art approaches for polymers conduct data augmentation to generate additional samples but unavoidably incurs extra computational costs. In contrast, large-scale open-source datasets are available for small molecules and provide a potential solution to data scarcity through transfer learning. In this work, we show that using transformers pretrained on small molecules and fine-tuned on polymer properties achieve comparable accuracy to those trained on augmented polymer datasets for a series of benchmark prediction tasks.

cs.LG

ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Deep Learning and High Performance Computing

Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models taken from NLP. These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive models (Transformer-XL, XLNet) and four auto-encoder models (BERT, Albert, Electra, T5) on data from UniRef and BFD containing up to 393 billion amino acids. The LMs were trained on the Summit supercomputer using 5616 GPUs and TPU Pod up-to 1024 cores. Dimensionality reduction revealed that the raw protein LM-embeddings from unlabeled data captured some biophysical features of protein sequences. We validated the advantage of using the embeddings as exclusive input for several subsequent tasks. The first was a per-residue prediction of protein secondary structure (3-state accuracy Q3=81%-87%); the second were per-protein predictions of protein sub-cellular localization (ten-state accuracy: Q10=81%) and membrane vs. water-soluble (2-state accuracy Q2=91%). For the per-residue predictions the transfer of the most informative embeddings (ProtT5) for the first time outperformed the state-of-the-art without using evolutionary information thereby bypassing expensive database searches. Taken together, the results implied that protein LMs learned some of the grammar of the language of life. To facilitate future work, we released our models at https://github.com/agemagician/ProtTrans.

cs.LG

DeepDriveMD: Deep-Learning Driven Adaptive Molecular Simulations for Protein Folding

Simulations of biological macromolecules play an important role in understanding the physical basis of a number of complex processes such as protein folding. Even with increasing computational power and evolution of specialized architectures, the ability to simulate protein folding at atomistic scales still remains challenging. This stems from the dual aspects of high dimensionality of protein conformational landscapes, and the inability of atomistic molecular dynamics (MD) simulations to sufficiently sample these landscapes to observe folding events. Machine learning/deep learning (ML/DL) techniques, when combined with atomistic MD simulations offer the opportunity to potentially overcome these limitations by: (1) effectively reducing the dimensionality of MD simulations to automatically build latent representations that correspond to biophysically relevant reaction coordinates (RCs), and (2) driving MD simulations to automatically sample potentially novel conformational states based on these RCs. We examine how coupling DL approaches with MD simulations can fold small proteins effectively on supercomputers. In particular, we study the computational costs and effectiveness of scaling DL-coupled MD workflows by folding two prototypical systems, viz., Fs-peptide and the fast-folding variant of the villin head piece protein. We demonstrate that a DL driven MD workflow is able to effectively learn latent representations and drive adaptive simulations. Compared to traditional MD-based approaches, our approach achieves an effective performance gain in sampling the folded states by at least 2.3x. Our study provides a quantitative basis to understand how DL driven MD simulations, can lead to effective performance gains and reduced times to solution on supercomputing resources.

cs.DC

Deep Generative Model Driven Protein Folding Simulation

Significant progress in computer hardware and software have enabled molecular dynamics (MD) simulations to model complex biological phenomena such as protein folding. However, enabling MD simulations to access biologically relevant timescales (e.g., beyond milliseconds) still remains challenging. These limitations include (1) quantifying which set of states have already been (sufficiently) sampled in an ensemble of MD runs, and (2) identifying novel states from which simulations can be initiated to sample rare events (e.g., sampling folding events). With the recent success of deep learning and artificial intelligence techniques in analyzing large datasets, we posit that these techniques can also be used to adaptively guide MD simulations to model such complex biological phenomena. Leveraging our recently developed unsupervised deep learning technique to cluster protein folding trajectories into partially folded intermediates, we build an iterative workflow that enables our generative model to be coupled with all-atom MD simulations to fold small protein systems on emerging high performance computing platforms. We demonstrate our approach in folding Fs-peptide and the $ββα$ (BBA) fold, FSD-EY. Our adaptive workflow enables us to achieve an overall root-mean squared deviation (RMSD) to the native state of 1.6$~Å$ and 4.4~$Å$ respectively for Fs-peptide and FSD-EY. We also highlight some emerging challenges in the context of designing scalable workflows when data intensive deep learning techniques are coupled to compute intensive MD simulations.

q-bio.BM

Deep diving into the comparative study of Choline dynamics using molecular dynamics simulation and neutron scattering technique

We present here the comparative study between the dynamics Choline and Tetra-methyl ammonium bromide. This is well known that deficiency in Choline would cause many severe diseases. No wonder why Choline is crucial component for our nutrients and dietary requirements. We present here a comprehensive study using all-atom molecular dynamics simulation combined with neutron scattering technique for solute behavior in aqueous solution. The solvent behavior is discussed in the follow up work

physics.chem-ph

A comparative study of solvent dynamics in Choline Bromide aqueous solution using combination of neutron scattering technique and molecular dynamics simulation

A comparative study between Choline and Tetra-methyl ammonium bromide is presented here. Choline which is a crucial component for our dietary requirements causes many diseases if there is deficiency. We used combined approach of all-atom molecular dynamics simulation coupled with neutron scattering technique to study mainly the solvent behavior in this study. There is follow up work where we discussed about the solute dynamical behavior.

physics.chem-ph

Effect of Nanodiamond surfaces on Drug Delivery Systems

The prospect of RNA nanotechnology is increasing because of its numerous potential applications especially in medical science. The spherical Nanodiamonds (NDs) are becoming popular because of their lesser toxicity, desirable mechanical, optical properties, functionality and available surface areas. On other hand RNAs are stable, flexible and easy to bind to the NDs. In this work, we have studied the tRNA dynamics on ND surface by high-resolution quasi-elastic neutron scattering spectroscopy and all atom molecular dynamics simulation technique to understand how the tRNA motion is affected by the presence of ND. The flexibly of the tRNA is analyzed by the Mean square displacement analysis that shows tRNA have a sharp increase around 230K in its hydrated form. The intermediate scattering function (ISF) representing the tRNA dynamics follows the logarithmic decay as proposed by the Mode Coupling theory (MCT). But most importantly the tRNA dynamics is found to be faster in presence of ND within 220K to 310K compared to the freestanding ones. This is as we have shown, is because of the swollen RNA molecule due the introduction of hydrophilic ND surface.

physics.bio-ph

Behavior of aqueous Tetrabutylammonium bromide - a combined approach of microscopic simulation and neutron scattering

Aqueous solution of tetrabutylammonium bromide is studied by quasi-elastic neutron scattering, to give information on the dynamic modes involving the ions present. Using a careful combination of two techniques, time-of-flight (TOF) and neutron spin echo (NSE), we de- couple the dynamic information in both the coherently and incoherently scattered signal from this system. We take advantage of the different intensity ratio of the two signals, as detected by each of the techniques, to achieve this decoupling. By using heavy water as the sol- vent, the tetrabutylammonium cation is the only hydrogen-containing species in the system and gives rise to a significant incoherent scattered intensity. The dynamic analysis of the incoherent signal (measured by TOF) leads to a translational diffusion coefficient of the cation as that is in good agreement with previous NMR, neutron scattering and tracer diffusion measurements. The dynamic analysis of the coherent signal observed at wave-vectors < 0.6 angstrom^(-1) (measured by NSE) leads to a significantly slower diffusive mode, by a factor of 2. This study explains that apparent difference and the reasons for that.

physics.chem-ph