Searcharxiv⌕ Search

arXiv · 2609.36885

RNA Design via Conditioned Flow Matching and Finite-Policy Reinforcement Learning

Abstract

RNA design aims to identify sequences that fold into specified secondary structures. Existing methods formulate the task as target-specific search or conditional generation. However, natural RNA evolution proceeds through sequence variation and selection, with compensatory substitutions, whereas these methods do not explicitly model this process. To address this limitation, we propose a two-stage framework comprising RNA Inverse-Folding Flow (RNA-IFlow) and RNA-IFlow-RL. RNA-IFlow uses structure-conditioned Dirichlet Flow Matching to model coordinated variation across the sequence, while RNA-IFlow-RL maps the learned flow to a pairing-preserving finite policy and refines it with thermodynamic feedback. Our framework achieves leading performance on multiple benchmarks, reaching 85.19% Pass@1 on Rfam-27. Further analyses reveal thermodynamic gains, policy dynamics, and robustness across settings. Our work couples coordinated variation with thermodynamic selection, offering a novel paradigm for RNA design.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zefeng Lin, Xianyong Fang, Tianfan Fu, Xiaohua Xu. 2026-09-29. RNA Design via Conditioned Flow Matching and Finite-Policy Reinforcement Learning. https://arxiv.org/abs/2609.36885

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Leveraging secondary-structure information for accurate nucleic acid structure prediction with OFoldNA

Recent advances in biomolecular structure prediction have enabled accurate modelling of increasingly complex molecular systems. However, nucleic acid structure prediction remains challenging because of conformational flexibility and the limited availability of high-quality 3D structural data. Secondary structure (SS) provides a more readily available layer of structural information that captures base-pairing relationships and folding topology. Here we present OFoldNA, an all-atom diffusion model that incorporates SS information into nucleic acid folding and protein--nucleic acid co-folding. Without external SS information, OFoldNA achieved leading performance on FoldBench for both nucleic acid monomer folding and protein--nucleic acid co-folding, with particularly strong performance on DNA monomers and protein--DNA interfaces involving longer nucleic acid chains. When accurate base-pairing information was provided, OFoldNA-SS2TS further improved both folding and co-folding accuracy, while partial SS information also yielded consistent gains. The same auxiliary branch can also be used for RNA SS prediction as OFoldNA-SS, which achieved the best out-of-distribution performance on CHANRG. Together, these results show that intermediate structural information such as nucleic acid SS can be leveraged to improve all-atom 3D modelling, providing a general direction for incorporating complementary structural modalities into molecular structure prediction and design.

q-bio.BM↗

How 'Foundational' Are Current Molecular Foundation Models?

Large-scale models have permeated the molecular sciences, yet what makes a model 'foundational' in this domain remains poorly defined. This paper proposes three testable criteria for assessing the foundational nature of molecular models: (i) generality across molecular entities, properties, and tasks; (ii) transferability to new applications with no or minimal task-specific retraining; and (iii) generalization beyond the training distribution. Applying these criteria to the state of the art reveals promising progress, particularly visible in biomolecular structure prediction and machine-learned interatomic potentials, although none of the approaches examined fully satisfies all three. Success is concentrated in domains where target properties are consistently defined and training data are abundant, with low label noise relative to physically meaningful variation. More broadly, progress in the molecular sciences appears to depend less on model scale alone than on the quality and structure of available data, as well as the incorporation of prior knowledge into models, prediction tasks, or downstream applications. This work shifts the notion of a molecular foundation model from a descriptive label to a testable hypothesis, offering a framework for assessing current models and guiding future developments.

q-bio.BM↗

Co-folding with a Soup of Representations

Co-folding models such as AlphaFold3, Protenix, ESMFold2, and OpenDDE have advanced rapidly, yet no single model consistently performs best across all biomolecular complexes. In this paper, we show that their pair representations encode complementary information that can be transferred across models to improve structure prediction. We introduce SoupFold, which combines pair representations from multiple co-folding models in a common representation space and generates structures from the combined representation. Importantly, SoupFold does not retrain the co-folding models and learns only simple mappings to transfer representations across models. We evaluate SoupFold on antibody-antigen, protein-protein, protein-ligand, molecular glue, GPCR, and oligomeric complex prediction using AlphaFold3, Protenix, ESMFold2, and OpenDDE. By combining representations across models, SoupFold improves over individual co-folding models across the considered benchmarks.

q-bio.BM↗