Searcharxiv⌕ Search

arXiv subjects

Giuseppe Sacco

Publications and source records attributed to Giuseppe Sacco.

2 recordsLinked to original sources

MERGE-RNA: a physics-based model to predict RNA secondary structure ensembles with chemical probing

RNA function is tied to secondary structure, operating through dynamic and heterogeneous structural ensembles. While current analysis tools typically output single static structures or averaged contact maps, chemical probing methods like DMS capture nucleotide-resolution signals representing the full structural ensemble, which remain difficult to interpret structurally. To address this, we present MERGE-RNA, a framework that describes and outputs RNA as a structural ensemble. By modeling the physics of the experimental pipeline, MERGE-RNA learns a small set of transferable and interpretable parameters, enabling the integration of measurements across different molecules, probe concentrations, and replicates in a single optimization to improve robustness. Our model employs a maximum-entropy principle to predict thermodynamic populations, with the minimal adjustments necessary to align the ensemble with experimental data. We validate MERGE-RNA on diverse RNAs, showing that it achieves structural accuracy surpassing standard pseudo-free-energy methods and yields ensembles better recapitulating measured DMS reactivity. Applied to the V. vulnificus adenine riboswitch, MERGE-RNA recovers the NMR-resolved conformations and their ligand-induced rearrangement, with population shifts matching the NMR-derived K_d. In a designed RNA construct for which we report new DMS data, MERGE-RNA deconvolves mixed states to reveal transient intermediate populations involved in strand displacement, dynamics invisible to methods based on enumerating a small number of structures.

q-bio.BM↗

Machine Learning for RNA Secondary Structure Prediction: a review of current methods and challenges

Predicting the secondary structure of RNA is a core challenge in computational biology, essential for understanding molecular function and designing novel therapeutics. The field has evolved from foundational but accuracy-limited thermodynamic approaches to a new data-driven paradigm dominated by machine learning and deep learning. These models learn folding patterns directly from data, leading to significant performance gains. This review surveys the modern landscape of these methods, covering single-sequence, evolutionary-based, and hybrid models that blend machine learning with biophysics. A central theme is the field's "generalization crisis," where powerful models were found to fail on new RNA families, prompting a community-wide shift to stricter, homology-aware benchmarking. In response to the underlying challenge of data scarcity, RNA foundation models have emerged, learning from massive, unlabeled sequence corpora to improve generalization. Finally, we look ahead to the next set of major hurdles-including the accurate prediction of complex motifs like pseudoknots, scaling to kilobase-length transcripts, incorporating the chemical diversity of modified nucleotides, and shifting the prediction target from static structures to the dynamic ensembles that better capture biological function. We also highlight the need for a standardized, prospective benchmarking system to ensure unbiased validation and accelerate progress.

q-bio.BM↗