SearcharxivSearch

arXiv subjects

S. Gnanakaran

Publications and source records attributed to S. Gnanakaran.

8 recordsLinked to original sources

Agentic AI platforms for autonomous training and rule induction of human-human and virus-human protein-protein interactions

We instruct an AI agent to construct two separate agentic AI platforms: one for autonomous training of predictive ML models for human-human and virus-human PPI, and the other for inducing explicit general rules governing human-human and virus-human PPI. The first agentic AI platform for autonomous training of predictive ML models for PPI is designed to consist of five AI agents that handle autonomous data collection, data verification, feature embedding, model design, and training and validation on three-way protein-disjoint cross-fold datasets. For human-human and human-virus PPIs, the final three-way protein-disjoint ensemble achieves an accuracy of 87.3% and 86.5%, respectively. For cross-checking and interpretability purposes, the second agentic AI platform is designed to replace ML predictions with human-readable rules derived from protein embeddings, physicochemical autocovariance descriptors, compartment annotations, pathway-domain overlap, and graph contexts. For human-human PPI, it is defined by a two-rule induction, whereas human-virus is induced by a more complex set of weighted rules. The rules induced by the second agentic platform align with the SHAP-identified features from the predictive ML models built by the first agentic platform. Taken together, our work demonstrates the agentic AI's ability to orchestrate from data planning to execution, and from rule induction to explanation in ML, opening the door to various applications.

cs.AI

Knowledge database development by large language models for countermeasures against viruses and marine toxins

Access to the most up-to-date information on medical countermeasures is important for the research and development of effective treatments for viruses and marine toxins. However, there is a lack of comprehensive databases that curate data on viruses and marine toxins, making decisions on medical countermeasures slow and difficult. In this work, we employ two large language models (LLMs) of ChatGPT and Grok to design two comprehensive databases of therapeutic countermeasures for five viruses of Lassa, Marburg, Ebola, Nipah, and Venezuelan equine encephalitis, as well as marine toxins. With high-level human-provided inputs, the two LLMs identify public databases containing data on the five viruses and marine toxins, collect relevant information from these databases and the literature, iteratively cross-validate the collected information, and design interactive webpages for easy access to the curated, comprehensive databases. Notably, the ChatGPT LLM is employed to design agentic AI workflows (consisting of two AI agents for research and decision-making) to rank countermeasures for viruses and marine toxins in the databases. Together, our work explores the potential of LLMs as a scalable, updatable approach for building comprehensive knowledge databases and supporting evidence-based decision-making.

cs.AI

Development of a Fragment-Based Machine Learning Algorithm for Designing Hybrid Drugs Optimized for Permeating Gram-Negative Bacteria

Gram-negative bacteria are a serious health concern due to the strong multidrug resistance that they display, partly due to the presence of a permeability barrier comprising two membranes with active efflux. New approaches are urgently needed to design antibiotics effective against these pathogens. In this work, we present a novel topological fragment-based approach ("Hunting Fragments Of X" or "Hunting FOX") to rationally "hunt for" chemical fragments that promote compound ability to permeate the outer membrane. Our approach generalizes to other drug design applications. We measure minimum inhibitory concentrations of compounds in two strains of Pseudomonas aeruginosa with variable permeability barriers and use them as an input to the Hunting FOX algorithm to identify molecular fragments responsible for enhanced outer membrane permeation properties and candidate molecules from an external library that demonstrate good permeation ability. Overall, we present proof of concept for a novel method that is expected to be valuable for rational design of hybrid drugs.

q-bio.QM

Computing Long Timescale Biomolecular Dynamics using Quasi-Stationary Distribution Kinetic Monte Carlo (QSD-KMC)

It is a challenge to obtain an accurate model of the state-to-state dynamics of a complex biological system from molecular dynamics (MD) simulations. In recent years, Markov State Models have gained immense popularity for computing state-to-state dynamics from a pool of short MD simulations. However, the assumption that the underlying dynamics on the reduced space is Markovian induces a systematic bias in the model, especially in biomolecular systems with complicated energy landscapes. To address this problem, we have devised a new approach we call quasi-stationary distribution kinetic Monte Carlo (QSD-KMC) that gives accurate long time state-to-state evolution while retaining the entire time resolution even when the dynamics is highly non-Markovian. The proposed method is a kinetic Monte Carlo approach that takes advantage of two concepts: (i) the quasi-stationary distribution and (ii) dynamical corrections theory. Implementation of QSD-KMC imposes stricter requirements on the lengths of the trajectories than in a Markov State Model approach, as the trajectories must be long enough to dephase. However, the QSD-KMC model produces state-to-state trajectories that are statistically indistinguishable from an MD trajectory mapped onto the discrete set of states, for an arbitrary choice of state decomposition. Furthermore, the aforementioned concepts can be used to construct a Monte Carlo approach to optimize the state boundaries regardless of the initial choice of states. We demonstrate the QSD-KMC method on two one-dimensional model systems, one of which is a driven nonequilibrium system, and on two well-characterized biomolecular systems.

physics.chem-ph

Dynamics of a Chlorophyll Dimer in Collective and Local Thermal Environments

We present a theoretical analysis of exciton transfer and decoherence effects in a photosynthetic dimer interacting with collective (correlated) and local (uncorrelated) protein-solvent environments. Our approach is based on the framework of the spin-boson model. We derive explicitly the thermal relaxation and decoherence rates of the exciton transfer process, valid for arbitrary temperatures and for arbitrary (in particular, large) interaction constants between the dimer and the environments. We establish a generalization of the Marcus formula, giving reaction rates for dimer levels possibly individually and asymmetrically coupled to environments. We identify rigorously parameter regimes for the validity of the generalized Marcus formula. The existence of long living quantum coherences at ambient temperatures emerges naturally from our approach.

quant-ph

Longitudinal Antigenic Sequences and Sites from Intra-Host Evolution (LASSIE) Identifies Immune-Selected HIV Variants

Within-host genetic sequencing from samples collected over time provides a dynamic view of how viruses evade host immunity. Immune-driven mutations might stimulate neutralization breadth by selecting antibodies adapted to cycles of immune escape that generate within-subject epitope diversity. Comprehensive identification of immune-escape mutations is experimentally and computationally challenging. With current technology, many more viral sequences can readily be obtained than can be tested for binding and neutralization, making down-selection necessary. Typically, this is done manually, by picking variants that represent different time-points and branches on a phylogenetic tree. Such strategies are likely to miss many relevant mutations and combinations of mutations, and to be redundant for other mutations. Longitudinal Antigenic Sequences and Sites from Intrahost Evolution (LASSIE) uses transmitted-founder loss to identify virus "hot-spots" under putative immune selection and chooses sequences that represent recurrent mutations in selected sites. LASSIE favors earliest sequences in which mutations arise. With well-characterized longitudinal Env sequences, we confirmed selected sites were concentrated in antibody contacts and selected sequences represented diverse antigenic phenotypes. Practical applications include rapidly identifying immune targets under selective pressure within a subject, selecting minimal sets of reagents for immunological assays that characterize evolving antibody responses, and for immunogens in polyvalent "cocktail" vaccines.

q-bio.PE

Sodium Chloride interaction with solvated and crystalline cellulose : sodium ion affects the tetramer and fibril in aqueous solution

Inorganic salts are a natural component of biomass which have a significant effect on the product yields from a variety of biomass conversion processes. Understanding their effect on biomass at the microscopic level can help discover their mechanistic role. We present a study of the effect of aqueous sodium chloride (NaCl) on the largest component of biomass, cellulose, focused on the thermodynamic and structural effect of a sodium ion on the cellulose tetramer, and fibril. Replica exchange molecular dynamics simulations of a cellulose tetramer reveal a number of preferred cellulose-Na contacts and bridging positions. Large scale MD simulations on a model cellulose fibril find that Na+ perturbs the hydroxymethyl rotational state population and consequently disrupts the "native" hydrogen bonding network.

physics.chem-ph

Network Structure of Protein Folding Pathways

The classical approach to protein folding inspired by statistical mechanics avoids the high dimensional structure of the conformation space by using effective coordinates. Here we introduce a network approach to capture the statistical properties of the structure of conformation spaces. Conformations are represented as nodes of the network, while links are transitions via elementary rotations around a chemical bond. Self-avoidance of a polypeptide chain introduces degree correlations in the conformation network, which in turn lead to energy landscape correlations. Folding can be interpreted as a biased random walk on the conformation network. We show that the folding pathways along energy gradients organize themselves into scale free networks, thus explaining previous observations made via molecular dynamics simulations. We also show that these energy landscape correlations are essential for recovering the observed connectivity exponent, which belongs to a different universality class than that of random energy models. In addition, we predict that the exponent and therefore the structure of the folding network fundamentally changes at high temperatures, as verified by our simulations on the AK peptide.

q-bio.BM