arXiv · 2412.12979
Reinforcement Learning Guides Generative Protein Language Models
Abstract
Protein engineering can optimize molecules for biotechnology and therapeutics, but navigating the high-dimensional sequence landscape remains challenging. Protein language models (pLMs) have shown to to generate functional proteins far from natural sequences, yet their outputs tend to reflect prevalent properties in training data, limiting discovery of rare properties such as high catalytic activity or thermostability. Here, we introduce ProtRL, a reinforcement learning framework for pLMs that iteratively updates model parameters to maximize externally defined reward functions. Across diverse design tasks, ProtRL shifts generation toward specified objectives while maintaining sequence diversity. We demonstrate the optimization of target folds, bounded and continuous fitness predictors, and multi-objective optimization in binder design. As a proof of concept, we applied ProtRL to experimental feedback for the engineering of epidermal growth factor receptor binders. Testing fewer than 100 designed variants across the experimental campaign, ProtRL-guided optimization provided a final round in which 16 of 22 variants bound EGFR. The best variant showed a dissociation constant of 5.5 nM, representing a nine-fold improvement over wild-type EGF and higher affinity than previously reported EGF variants identified through substantially larger screening campaigns. Our code and models are publicly available at github.com/AI4PDLab/ProtRL
Explore related subjects
Keep this discovery
Filippo Stocco, Maria Artigues-Lleixa, Andrea Hunklinger, Michele Garibbo, Talal Widatalla, Marc Guell, Noelia Ferruz. 2024-12-17. Reinforcement Learning Guides Generative Protein Language Models. https://arxiv.org/abs/2412.12979
Cite the original work for its findings. Save a collection to share your selection of sources.