SearcharxivSearch

arXiv subjects

Noelia Ferruz

Publications and source records attributed to Noelia Ferruz.

6 recordsLinked to original sources

Reinforcement Learning Guides Generative Protein Language Models

Protein engineering can optimize molecules for biotechnology and therapeutics, but navigating the high-dimensional sequence landscape remains challenging. Protein language models (pLMs) have shown to to generate functional proteins far from natural sequences, yet their outputs tend to reflect prevalent properties in training data, limiting discovery of rare properties such as high catalytic activity or thermostability. Here, we introduce ProtRL, a reinforcement learning framework for pLMs that iteratively updates model parameters to maximize externally defined reward functions. Across diverse design tasks, ProtRL shifts generation toward specified objectives while maintaining sequence diversity. We demonstrate the optimization of target folds, bounded and continuous fitness predictors, and multi-objective optimization in binder design. As a proof of concept, we applied ProtRL to experimental feedback for the engineering of epidermal growth factor receptor binders. Testing fewer than 100 designed variants across the experimental campaign, ProtRL-guided optimization provided a final round in which 16 of 22 variants bound EGFR. The best variant showed a dissociation constant of 5.5 nM, representing a nine-fold improvement over wild-type EGF and higher affinity than previously reported EGF variants identified through substantially larger screening campaigns. Our code and models are publicly available at github.com/AI4PDLab/ProtRL

q-bio.BM

Steering Generative Models for Protein Design: Aligning and Conditioning Strategies

Generative artificial intelligence models learn probability distributions from data and produce novel samples that capture the salient properties of their training sets. Proteins are particularly attractive for such approaches given their abundant data and the versatility of their representations, ranging from sequences to structures and functions. This versatility has motivated the rapid development of generative models for protein design, enabling the generation of functional proteins and enzymes with unprecedented success. However, because these models mirror their training distribution, they tend to sample from its most probable modes, while low-probability regions, often encoding valuable properties, remain underexplored. To address this challenge, recent work has proposed strategies for steering generative models toward user-specified properties. In this review, we survey and categorize these strategies, distinguishing approaches that modify model parameters, such as reinforcement learning or supervised fine-tuning, from those that keep the model's parameters fixed, including conditional generation, retrieval-augmented strategies, Bayesian guidance, and tailored sampling methods. Together, these developments are beginning to enable the steering of generative models toward proteins with desired properties.

q-bio.BM

Generative AI for Enzyme Design and Biocatalysis

Sparked by innovations in generative artificial intelligence (AI), the field of protein design has undergone a paradigm shift with an explosion of new models for optimizing existing enzymes or creating them from scratch. After more than one decade of low success rates for computationally designed enzymes, generative AI models are now frequently used for designing proficient enzymes. Here, we provide a comprehensive overview and classification of generative AI models for enzyme design, highlighting models with experimental validation relevant to real-world settings and outlining their respective limitations. We argue that generative AI models now have the maturity to create and optimize enzymes for industrial applications. Wider adoption of generative AI models with experimental feedback loops can speed up the development of biocatalysts and serve as a community assessment to inform the next generation of models.

q-bio.BM

Toward the Explainability of Protein Language Models

Protein language models (pLMs) excel in a variety of tasks that range from structure prediction to the design of functional enzymes. However, these models operate as black boxes, and their underlying working principles remain unclear. Here, we survey emerging applications of explainable artificial intelligence (XAI) to pLMs and describe the potential of XAI in protein research. We divide the workflow of protein AI modeling into four information contexts: (i) training sequences, (ii) input prompt, (iii) model architecture, and (iv) input-output pairs. For each, we describe existing methods and applications of XAI. Additionally, from published studies we distil five (potential) roles that XAI can play in protein research: Evaluator, Multitasker, Engineer, Coach, and Teacher, with the Evaluator role being the only one widely adopted so far. These roles aim to help both protein scientists and model developers understand the possibilities and limitations of implementing XAI for predictive and generative tasks. While our analysis focuses on pLMs, both this categorization and roles are broadly applicable to any other model architectures. We conclude by highlighting critical areas of application for the future, including risks related to security, trustworthiness, and bias, and we call for community benchmarks, open-source tooling, domain-specific visualizations, and wet-lab characterization to advance the interpretability of protein AI.

q-bio.BM

Exploring the Protein Sequence Space with Global Generative Models

Recent advancements in specialized large-scale architectures for training image and language have profoundly impacted the field of computer vision and natural language processing (NLP). Language models, such as the recent ChatGPT and GPT4 have demonstrated exceptional capabilities in processing, translating, and generating human languages. These breakthroughs have also been reflected in protein research, leading to the rapid development of numerous new methods in a short time, with unprecedented performance. Language models, in particular, have seen widespread use in protein research, as they have been utilized to embed proteins, generate novel ones, and predict tertiary structures. In this book chapter, we provide an overview of the use of protein generative models, reviewing 1) language models for the design of novel artificial proteins, 2) works that use non-Transformer architectures, and 3) applications in directed evolution approaches.

q-bio.BM

Controllable Protein Design with Language Models

The 21st century is presenting humankind with unprecedented environmental and medical challenges. The ability to design novel proteins tailored for specific purposes could transform our ability to respond timely to these issues. Recent advances in the field of artificial intelligence are now setting the stage to make this goal achievable. Protein sequences are inherently similar to natural languages: Amino acids arrange in a multitude of combinations to form structures that carry function, the same way as letters form words and sentences that carry meaning. Therefore, it is not surprising that throughout the history of Natural Language Processing (NLP), many of its techniques have been applied to protein research problems. In the last few years, we have witnessed revolutionary breakthroughs in the field of NLP. The implementation of Transformer pre-trained models has enabled text generation with human-like capabilities, including texts with specific properties such as style or subject. Motivated by its considerable success in NLP tasks, we expect dedicated Transformers to dominate custom protein sequence generation in the near future. Finetuning pre-trained models on protein families will enable the extension of their repertoires with novel sequences that could be highly divergent but still potentially functional. The combination of control tags such as cellular compartment or function will further enable the controllable design of novel protein functions. Moreover, recent model interpretability methods will allow us to open the 'black box' and thus enhance our understanding of folding principles. While early initiatives show the enormous potential of generative language models to design functional sequences, the field is still in its infancy. We believe that protein language models are a promising and largely unexplored field and discuss their foreseeable impact on protein design.

q-bio.BM