Searcharxiv⌕ Search

arXiv · 2609.34001

Assay-Aware BindingDB: Curating Experimental Context for Binding Affinity Prediction

Abstract

Protein--ligand binding affinity prediction is fundamental to computational drug discovery, yet modern AI-driven models are limited by pervasive heterogeneity in their training data: bioactivity values are aggregated across diverse assay types and experimental conditions without accounting for protocol-level differences, introducing systematic noise. Existing harmonization approaches either discard assay-level metadata or collapse it into coarse categorical distinctions, leaving rich contextual signal unused. We address this gap with two contributions. First, we introduce Assay-Aware BindingDB, augmenting 74,425 BindingDB protein--ligand pairs across four assay types (ITC, SPR, RBA, and FPA) with structured metadata extracted from primary literature using a two-stage agentic framework that separates evidence extraction from ontology-conditioned JSON synthesis. Against domain-expert references, the framework achieves presence F1 $\geq 0.947$ and semantic content accuracy $\geq 0.913$ across all assay types. Second, we develop a context-conditioned affinity model that injects an assay-context embedding into the Boltz-2 affinity module. On a paper-level held-out split, the model reduces combined-regime MSE from $1.27$ to $1.19$ and increases Pearson correlation from $0.64$ to $0.67$, with significant gains for SPR and FPA. ITC, the only label- and immobilization-free assay considered, shows no improvement, consistent with its design removing the protocol artifacts the metadata captures. These results support the hypothesis that systematically curated assay metadata provides informative signal for affinity prediction.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ming-Hsiu Wu, Xuejiao Shirley Guo, Bingsong Zeng, Ziqian Xie, Shuiwang Ji, Wenshe Liu, Cui Tao, Degui Zhi. 2026-09-27. Assay-Aware BindingDB: Curating Experimental Context for Binding Affinity Prediction. https://arxiv.org/abs/2609.34001

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Co-folding with a Soup of Representations

Co-folding models such as AlphaFold3, Protenix, ESMFold2, and OpenDDE have advanced rapidly, yet no single model consistently performs best across all biomolecular complexes. In this paper, we show that their pair representations encode complementary information that can be transferred across models to improve structure prediction. We introduce SoupFold, which combines pair representations from multiple co-folding models in a common representation space and generates structures from the combined representation. Importantly, SoupFold does not retrain the co-folding models and learns only simple mappings to transfer representations across models. We evaluate SoupFold on antibody-antigen, protein-protein, protein-ligand, molecular glue, GPCR, and oligomeric complex prediction using AlphaFold3, Protenix, ESMFold2, and OpenDDE. By combining representations across models, SoupFold improves over individual co-folding models across the considered benchmarks.

q-bio.BM↗

Continuous Variational Synthesis

Biological machine learning was long bottlenecked by the ability to synthesize designed DNA. Variational synthesis models control chemical reactions to physically manufacture quadrillions of designed sequences in DNA. However, training these generative models is challenging: constraints on chemical synthesis can force many parameters into a discrete space, limiting the ability to pre-train and fine-tune. In this article we train ``free'' variational synthesis models using stochastic gradient descent in continuous space, and then discretize with post-training quantization to impose hardware and wetware constraints. This enables variational synthesis models to satisfy stringent reward criteria, while still synthesizing diverse designs, achieving a strictly dominating quality-diversity Pareto frontier. We demonstrate by training variational synthesis models of enzymes, peptides, antibody CDRH3s, and regulatory DNA elements. In silico performance is maintained in vitro.

q-bio.BM↗

Where Should Physics Enter a Molecular Crystal Generator?

Generative models make molecular crystal structure prediction fast, but their samples still exhibit geometric and packing violations. Physics can be introduced during training, post-training, or inference, yet these choices are rarely compared with the generator and physical signal held fixed. We introduce CrystAF, an all-atom crystal flow-map generation model, and use it with the UMA interatomic potential to systematically study where physics should enter. Post-training learns physical preferences directly into CrystAF, improving molecular validity and crystal packing while leaving sampling unchanged: physics is paid for once during training rather than repeatedly at deployment. In contrast, UMA relaxation is effective at repairing local clashes but makes generation 6--26$\times$ slower, while learning from relaxed targets provides little benefit. These routes are complementary rather than competing. Physics-informed post-training first shifts the generated distribution toward more physically reasonable structures, after which inexpensive inference-time corrections further remove clashes and restore stereochemistry that the generator cannot represent. Importantly, the same post-training strategy also improves the multi-step all-atom Clari-M and rigid-body MolCrystalFlow generators, demonstrating transfer across architectures and representations. Together, our results suggest a simple principle: learn reusable physical alignment into the generator, and reserve inference-time physics for residual constraints that are better corrected than learned.

q-bio.BM↗