SearcharxivSearch

arXiv subjects

Sohvi Luukkonen

Publications and source records attributed to Sohvi Luukkonen.

7 recordsLinked to original sources

Contrastive Geometric Learning Unlocks Unified Structure- and Ligand-Based Drug Design

Structure-based and ligand-based computational drug design have traditionally relied on disjoint data sources and modeling assumptions, limiting their joint use at scale. In this work, we introduce Contrastive Geometric Learning for Unified Computational Drug Design (ConGLUDe), a single contrastive geometric model that unifies structure- and ligand-based training. ConGLUDe couples a geometric protein encoder that produces whole-protein representations and implicit embeddings of predicted binding sites with a fast ligand encoder, removing the need for predefined pockets. By aligning ligands with both global protein representations and multiple candidate binding sites through contrastive learning, ConGLUDe supports ligand-conditioned pocket prediction in addition to virtual screening and target fishing, while being trained jointly on protein-ligand complexes and large-scale bioactivity data. Across diverse benchmarks, ConGLUDe achieves competitive zero-shot virtual screening performance, substantially outperforms existing methods on a challenging target fishing task, and demonstrates state-of-the-art ligand-conditioned pocket selection. These results highlight the advantages of unified structure-ligand training and position ConGLUDe as a step toward general-purpose foundation models for drug discovery.

cs.LG

MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs

A molecule's properties are fundamentally determined by its composition and structure encoded in its molecular graph. Thus, reasoning about molecular properties requires the ability to parse and understand the molecular graph. Large Language Models (LLMs) are increasingly applied to chemistry, tackling tasks such as molecular name conversion, captioning, text-guided generation, and property or reaction prediction. Most existing benchmarks emphasize general chemical knowledge, rely on literature or surrogate labels that risk leakage or bias, or reduce evaluation to multiple-choice questions. We introduce MolecularIQ, a molecular structure reasoning benchmark focused exclusively on symbolically verifiable tasks. MolecularIQ enables fine-grained evaluation of reasoning over molecular graphs and reveals capability patterns that localize model failures to specific tasks and molecular structures. This provides actionable insights into the strengths and limitations of current chemistry LLMs and guides the development of models that reason faithfully over molecular structure.

cs.LG

Measuring AI Progress in Drug Discovery: A Reproducible Leaderboard for the Tox21 Challenge

Deep learning's rise since the early 2010s has transformed fields like computer vision and natural language processing and strongly influenced biomedical research. For drug discovery specifically, a key inflection - akin to vision's "ImageNet moment" - arrived in 2015, when deep neural networks surpassed traditional approaches on the Tox21 Data Challenge. This milestone accelerated the adoption of deep learning across the pharmaceutical industry, and today most major companies have integrated these methods into their research pipelines. After the Tox21 Challenge concluded, its dataset was included in several established benchmarks, such as MoleculeNet and the Open Graph Benchmark. However, during these integrations, the dataset was altered and labels were imputed or manufactured, resulting in a loss of comparability across studies. Consequently, the extent to which bioactivity and toxicity prediction methods have improved over the past decade remains unclear. To this end, we introduce a reproducible leaderboard, hosted on Hugging Face with the original Tox21 Challenge dataset, together with a set of baseline and representative methods. The current version of the leaderboard indicates that the original Tox21 winner - the ensemble-based DeepTox method - and the descriptor-based self-normalizing neural networks introduced in 2017, continue to perform competitively and rank among the top methods for toxicity prediction, leaving it unclear whether substantial progress in toxicity prediction has been achieved over the past decade. As part of this work, we make all baselines and evaluated models publicly accessible for inference via standardized API calls to Hugging Face Spaces.

cs.LG

Bio-xLSTM: Generative modeling, representation and in-context learning of biological and chemical sequences

Language models for biological and chemical sequences enable crucial applications such as drug discovery, protein engineering, and precision medicine. Currently, these language models are predominantly based on Transformer architectures. While Transformers have yielded impressive results, their quadratic runtime dependency on the sequence length complicates their use for long genomic sequences and in-context learning on proteins and chemical sequences. Recently, the recurrent xLSTM architecture has been shown to perform favorably compared to Transformers and modern state-space model (SSM) architectures in the natural language domain. Similar to SSMs, xLSTMs have a linear runtime dependency on the sequence length and allow for constant-memory decoding at inference time, which makes them prime candidates for modeling long-range dependencies in biological and chemical sequences. In this work, we tailor xLSTM towards these domains and propose a suite of architectural variants called Bio-xLSTM. Extensive experiments in three large domains, genomics, proteins, and chemistry, were performed to assess xLSTM's ability to model biological and chemical sequences. The results show that models based on Bio-xLSTM a) can serve as proficient generative models for DNA, protein, and chemical sequences, b) learn rich representations for those modalities, and c) can perform in-context learning for proteins and small molecules.

q-bio.BM

Pressure Correction for Solvation Theories

Liquid state theories such as integral equations and classical density functional theory often overestimate the bulk pressure of fluids because they require closure relations or truncations of functionals. Consequently, the cost to create a molecular cavity in the fluid is no longer negligible and those theories predict wrong solvation free energies. We show how to correct them simply by computing an optimized Van der Walls volume of the solute and removing the undue free energy to create such volume in the fluid. Given this versatile correction, we demonstrate that state-of-the-art solvation theories can predict, within seconds, hydration free energies of a benchmark of small neutral drug-like molecules with the same accuracy as day-long molecular simulations.

physics.chem-ph

Predicting hydration free energies of the FreeSolv database of druglike molecules with molecular density functional theory

We assess the performance of molecular densityfunctional theory (MDFT) to predict hydration freeenergies of the small drug-like molecules benchmark,FreeSolv. MDFT in the hyper-netted chain approx-imation (HNC) coupled with a pressure correctionpredicts experimental hydration free energies of theFreeSolv database within 1 kcal/mol with an averagecomputation time of two cpu.min per molecule. Thisis the same accuracy as for simulation based free en-ergy calculations that typically require hundreds ofcpu.h or tens of gpu.h per molecule.

physics.chem-ph

High-throughput free energies and water maps for drug discovery by molecular density functional theory

The hydration or binding free energy of a drug-like molecule is a key data for early stage drug discovery. Hundreds of thousands of evaluations are needed, which rules out the exhaustive use of atomistic simulations and free energy methods. Instead, the current docking and screening processes are today relying on numerically efficient scoring functions that lose much of the atomic scale information and hence remain error-prone. In this article, we show how a probabilistic description of molecular liquids as implemented in the molecular density functional theory predicts hydration free energies of a state-of-the-art benchmark of small drug-like molecules within 0.5 kJ/mol (0.1 kcal/mol) of atomistic simulations, along with water and polarization maps, for a computation time compatible with screening and docking.

physics.chem-ph