Searcharxiv⌕ Search

arXiv subjects

Srimukh Prasad Veccham

Publications and source records attributed to Srimukh Prasad Veccham.

11 recordsLinked to original sources

Probabilistic Contrastive Pretraining for Multi-task ADME Property Prediction

Accurate prediction of absorption, distribution, metabolism, and excretion (ADME) properties is critical to drug discovery, but remains challenging because ADME endpoints are noisy, interdependent, and often data-limited. We propose a molecular graph-transformer pretraining framework that combines chemistry-specific self-supervision with contrastive mutual information machine learning (cMIM). Our method encodes molecular graphs into latent variables, reconstructs SMILES strings from the graph-derived latent codes, and augments the contrastive objective with domain-specific self-supervised chemistry tasks. Rather than treating these tasks as auxiliary regularizers with separately tuned loss weights, we formulate reconstruction, contrastive discrimination, and chemistry-specific supervision as unit-weighted log-probability factors in a single probabilistic latent-variable objective. For fine-tuning, we propose a multi-task GNN readout architecture with task-specific multilayer perceptron heads, preserving shared representation learning while mitigating negative transfer and improving the modeling of heterogeneous, nonlinear task relationships. Across Biogen, ExpansionRX, and ChEMBL-MT, the resulting Contrastive KERMT pretraining improves over the KERMT baseline by 7.6%, 9.9%, and 9.5% respectively (averaged over significantly-improved endpoints). Adding ADME-adjacent molecules to the pretraining corpus further improves transfer, and the contrastive component sharpens chemically meaningful latent neighborhoods.

cs.LG↗

Exploring Synthesizable Chemical Space with Iterative Pathway Refinements

A well-known pitfall of molecular generative models is that they are not guaranteed to generate synthesizable molecules. Existing solutions for this problem often struggle to effectively navigate exponentially large combinatorial space of synthesizable molecules and suffer from poor coverage. To address this problem, we introduce ReaSyn, an iterative generative pathway refinement framework that obtains synthesizable analogs to input molecules by projecting them onto synthesizable space. Specifically, we propose a simple synthetic pathway representation that allows for generating pathways in both bottom-up and top-down traversal of synthetic trees. We design ReaSyn so that both bottom-up and top-down pathways can be sampled with a single unified autoregressive model. ReaSyn can thus iteratively refine subtrees of generated synthetic trees in a bidirectional manner. Further, we introduce a discrete flow model that refines the generated pathway at the entire pathway level with edit operations: insertion, deletion, and substitution. The iterative refinement cycle of (1) bottom-up decoding, (2) top-down decoding, and (3) holistic editing constitutes a powerful pathway reasoning strategy, allowing the model to explore the vast space of synthesizable molecules. Experimentally, ReaSyn achieves the highest reconstruction rate and pathway diversity in synthesizable molecule reconstruction and the highest optimization performance in synthesizable goal-directed molecular optimization, and significantly outperforms previous synthesizable projection methods in synthesizable hit expansion. These results highlight ReaSyn's superior ability to navigate combinatorially-large synthesizable chemical space.

cs.LG↗

Multitask finetuning and acceleration of chemical pretrained models for small molecule drug property prediction

Chemical pretrained models, sometimes referred to as foundation models, are receiving considerable interest for drug discovery applications. The general chemical knowledge extracted from self-supervised training has the potential to improve predictions for critical drug discovery endpoints, including on-target potency and ADMET properties. Multi-task learning has previously been successfully leveraged to improve predictive models. Here, we show that enabling multitasking in finetuning of chemical pretrained graph neural network models such as Kinetic GROVER Multi-Task (KERMT), an enhanced version of the GROVER model, and Knowledge-guided Pre-training of Graph Transformer (KGPT) significantly improves performance over non-pretrained graph neural network models. Surprisingly, we find that the performance improvement from finetuning KERMT in a multitask manner is most significant at larger data sizes. Additionally, we publish two multitask ADMET data splits to enable more accurate benchmarking of multitask deep learning methods for drug property prediction. Finally, we provide an accelerated implementation of the KERMT model on GitHub, unlocking large-scale pretraining, finetuning, and inference in industrial drug discovery workflows.

cs.LG↗

Want to train KANS at scale? Now UKAN!

Kolmogorov-Arnold Networks (KANs) have recently emerged as a powerful alternative to traditional multilayer perceptrons. However, their reliance on predefined, bounded grids restricts their ability to approximate functions on unbounded domains. To address this, we present Unbounded Kolmogorov-Arnold Networks (UKANs), a method that removes the need for bounded grids in traditional Kolmogorov-Arnold Networks (KANs). The key innovation of this method is a coefficient-generator (CG) model that produces, on the fly, only the B-spline coefficients required locally on an unbounded symmetric grid. UKANs couple multilayer perceptrons with KANs by feeding the positional encoding of grid groups into the CG model, enabling function approximation on unbounded domains without requiring data normalization. To reduce the computational cost of both UKANs and KANs, we introduce a GPU-accelerated library that lowers B-spline evaluation complexity by a factor proportional to the grid size, enabling large-scale learning by leveraging efficient memory management, in line with recent software advances such as FlashAttention and FlashFFTConv. Performance benchmarking confirms the superior memory and computational efficiency of our accelerated KAN (warpKAN), and UKANs, showing a 3-30x speed-up and up to 1000x memory reduction compared to vanilla KANs. Experiments on regression, classification, and generative tasks demonstrate the effectiveness of UKANs to match or surpass KAN accuracy. Finally, we use both accelerated KAN and UKAN in a molecular property prediction task, establishing the feasibility of large-scale end-to-end training with our optimized implementation.

cs.LG↗

BioNeMo Framework: a modular, high-performance library for AI model development in drug discovery

Artificial Intelligence models encoding biology and chemistry are opening new routes to high-throughput and high-quality in-silico drug development. However, their training increasingly relies on computational scale, with recent protein language models (pLM) training on hundreds of graphical processing units (GPUs). We introduce the BioNeMo Framework to facilitate the training of computational biology and chemistry AI models across hundreds of GPUs. Its modular design allows the integration of individual components, such as data loaders, into existing workflows and is open to community contributions. We detail technical features of the BioNeMo Framework through use cases such as pLM pre-training and fine-tuning. On 256 NVIDIA A100s, BioNeMo Framework trains a three billion parameter BERT-based pLM on over one trillion tokens in 4.2 days. The BioNeMo Framework is open-source and free for everyone to use.

cs.LG↗

GenMol: A Drug Discovery Generalist with Discrete Diffusion

Drug discovery is a complex process that involves multiple stages and tasks. However, existing molecular generative models can only tackle some of these tasks. We present Generalist Molecular generative model (GenMol), a versatile framework that uses only a single discrete diffusion model to handle diverse drug discovery scenarios. GenMol generates Sequential Attachment-based Fragment Embedding (SAFE) sequences through non-autoregressive bidirectional parallel decoding, thereby allowing the utilization of a molecular context that does not rely on the specific token ordering while having better sampling efficiency. GenMol uses fragments as basic building blocks for molecules and introduces fragment remasking, a strategy that optimizes molecules by regenerating masked fragments, enabling effective exploration of chemical space. We further propose molecular context guidance (MCG), a guidance method tailored for masked discrete diffusion of GenMol. GenMol significantly outperforms the previous GPT-based model in de novo generation and fragment-constrained generation, and achieves state-of-the-art performance in goal-directed hit generation and lead optimization. These results demonstrate that GenMol can tackle a wide range of drug discovery tasks, providing a unified and versatile approach for molecular design. Our code is available at https://github.com/NVIDIA-Digital-Bio/genmol.

cs.LG↗

Molecule Generation with Fragment Retrieval Augmentation

Fragment-based drug discovery, in which molecular fragments are assembled into new molecules with desirable biochemical properties, has achieved great success. However, many fragment-based molecule generation methods show limited exploration beyond the existing fragments in the database as they only reassemble or slightly modify the given ones. To tackle this problem, we propose a new fragment-based molecule generation framework with retrieval augmentation, namely Fragment Retrieval-Augmented Generation (f-RAG). f-RAG is based on a pre-trained molecular generative model that proposes additional fragments from input fragments to complete and generate a new molecule. Given a fragment vocabulary, f-RAG retrieves two types of fragments: (1) hard fragments, which serve as building blocks that will be explicitly included in the newly generated molecule, and (2) soft fragments, which serve as reference to guide the generation of new fragments through a trainable fragment injection module. To extrapolate beyond the existing fragments, f-RAG updates the fragment vocabulary with generated fragments via an iterative refinement process which is further enhanced with post-hoc genetic fragment modification. f-RAG can achieve an improved exploration-exploitation trade-off by maintaining a pool of fragments and expanding it with novel and high-quality fragments through a strong generative prior.

cs.LG↗

Assessment of the Performance of Density Functionals for Predicting Potential Energy Curves in Hydrogen Storage Applications

The availability of accurate computational tools for modeling and simulation is vital to accelerate the discovery of materials capable of storing hydrogen (H2) under given parameters of pressure swing and temperature. Previously, we compiled the H2Bind275 dataset consisting of equilibrium geometries and assessed the performance of 55 density functionals over this dataset (Veccham, S. P.; Head-Gordon, M. J. Chem. Theory Comput., 2020, 16, 4963--4982). As it is crucial for computational tools to accurately model the entire potential energy curve (PEC), in addition to the equilibrium geometry, we have extended this dataset with 389 new data points to include two compressed and three elongated geometries along 78 PECs for H2 binding, forming the H2Bind78x7 dataset. Assessing the performance of 55 density functionals on this significantly larger and more comprehensive H2Bind78x7 dataset, we have identified the best performing density functionals for H2 binding applications: PBE0-DH, $ω$B97X-V, $ω$B97M-V, and DSD-PBEPBE-D3(BJ). Addition of Hartree Fock exchange improves the performance of density functionals, albeit not uniformly throughout the PEC. We recommend the usage of wB97X-V and wB97M-V density functionals as they give good performance for both geometries and energies. In addition, we have also identified B97M-V and B97M-rV as the best semi-local density functionals for predicting H2 binding energy at its equilibrium geometry.

physics.chem-ph↗

A Non-Perturbative Pairwise-Additive Analysis of Charge Transfer Contributions to Intermolecular Interaction Energies

Energy decomposition analysis (EDA) based on absolutely localized molecular orbitals (ALMOs) decomposes the interaction energy between molecules into physically interpretable components like geometry distortion, frozen interactions, polarization, and charge transfer (CT, also sometimes called charge delocalization) interactions. In this work, a numerically exact scheme to decompose the CT interaction energy into pairwise additive terms is introduced for the ALMO-EDA using density functional theory. Unlike perturbative pairwise charge-decomposition analysis, the new approach does not break down for strongly interacting systems, or show significant exchange-correlation functional dependence in the decomposed energy components. Both the energy lowering and the charge flow associated with CT can be decomposed. Complementary occupied-virtual orbital pairs (COVPs) that capture the dominant donor and acceptor CT orbitals are obtained for the new decomposition. It is applied to systems with different types of interactions including DNA base-pairs, borane-ammonia adducts, and transition metal hexacarbonyls. While consistent with most existing understanding of the nature of CT in these systems, the results also reveal some new insights into the origin of trends in donor-acceptor interactions.

physics.chem-ph↗

Density Functionals for Hydrogen Storage: Defining the H2Bind275 Test Set with Ab Initio Benchmarks and Assessment of 55 Functionals

Efficient and high capacity storage materials are indispensable for a hydrogen-based economy. In silico tools can accelerate the process of discovery of new adsorbent materials with optimal hydrogen adsorption enthalpies. Density functional theory is well-poised to become a very useful tool for enabling high-throughput screening of potential materials. In this work, we have identified density functional approximations that provide good performance for hydrogen binding applications following a two-pronged approach. First, we have compiled a dataset (H2Bind275) that comprehensively represents the hydrogen binding problem capturing the chemical and mechanistic diversity in the binding sites encountered in hydrogen storage materials. We have also computed reference interaction energies for this dataset using coupled cluster theory. Secondly, we have assessed the performance of 55 density functional approximations for predicting H$_2$ interaction energies and have identified two hybrid density functionals ($ω$B97X-V and $ω$B97M-V), two double hybrid density functionals (DSD-PBEPBE-D3(BJ) and PBE0-DH), and one semi-local density functional (B97M-V) as the best performing ones. We have recommended the addition of empirical dispersion corrections to systematically underbinding density functionals like revPBE, BLYP, and B3LYP for improvements in performance at negligible additional cost. We have also recommended the usage of the def2-TZVPP basis set as it represents a good compromise between accuracy and cost, limiting the finite basis set errors to less than 1kJ/mol.

physics.chem-ph↗

Making Many-Body Interactions Nearly Pairwise Additive: The Polarized Many-Body Expansion Approach

The Many-Body Expansion (MBE) is a useful tool to simulate condensed phase chemical systems, often avoiding the steep computational cost of usual electronic structure methods. However, it often requires higher than 2-body terms to achieve quantitative accuracy. In this work, we propose the Polarized MBE (PolBE) method where each MBE energy contribution is treated as an embedding problem. In each energy term, a smaller fragment is embedded into a larger, polarized environment and only a small region is treated at the high-level of theory using embedded mean-field theory. The role of polarized environment was found to be crucial in providing quantitative accuracy at the 2-body level. PolBE accurately predicts non-covalent interaction energies for a number of systems, including CO$_2$, water, and hydrated ion clusters, with a variety of interaction mechanisms, from weak dispersion to strong electrostatics considered in this work. We further demonstrate that the PolBE interaction energy is predominantly pairwise unlike the usual vacuum MBE which requires higher-order terms to achieve similar accuracy. We numerically show that PolBE often performs better than other widely used embedded MBE methods such as the electrostatically embedded MBE. Owing to the lack of expensive diagonalization of Fock matrices and its embarrassingly parallel nature, PolBE is a promising way to access condensed phase systems with hybrid density functionals that are difficult to treat with currently available methods.

physics.chem-ph↗