SearcharxivSearch

arXiv subjects

Hongbin Shen

Publications and source records attributed to Hongbin Shen.

2 recordsLinked to original sources

Modeling enzyme temperature stability from sequence segment perspective

Developing enzymes with desired thermal properties is crucial for a wide range of industrial and research applications, and determining temperature stability is an essential step in this process. Experimental determination of thermal parameters is labor-intensive, time-consuming, and costly. Moreover, existing computational approaches are often hindered by limited data availability and imbalanced distributions. To address these challenges, we introduce a curated temperature stability dataset designed for model development and benchmarking in enzyme thermal modeling. Leveraging this dataset, we present the \textit{Segment Transformer}, a novel deep learning framework that enables efficient and accurate prediction of enzyme temperature stability. The model achieves state-of-the-art performance with an RMSE of 24.03, MAE of 18.09, and Pearson and Spearman correlations of 0.33, respectively. These results highlight the effectiveness of incorporating segment-level representations, grounded in the biological observation that different regions of a protein sequence contribute unequally to thermal behavior. As a proof of concept, we applied the Segment Transformer to guide the engineering of a cutinase enzyme. Experimental validation demonstrated a 1.64-fold improvement in relative activity following heat treatment, achieved through only 17 mutations and without compromising catalytic function.

cs.LG

GenoHoption: Bridging Gene Network Graphs and Single-Cell Foundation Models

The remarkable success of foundation models has sparked growing interest in their application to single-cell biology. Models like Geneformer and scGPT promise to serve as versatile tools in this specialized field. However, representing a cell as a sequence of genes remains an open question since the order of genes is interchangeable. Injecting the gene network graph offers gene relative positions and compact data representation but poses a dilemma: limited receptive fields without in-layer message passing or parameter explosion with message passing in each layer. To pave the way forward, we propose GenoHoption, a new computational framework for single-cell sequencing data that effortlessly combines the strengths of these foundation models with explicit relationships in gene networks. We also introduce a constraint that lightens the model by focusing on learning the predefined graph structure while ensuring further hops are deducted to expand the receptive field. Empirical studies show that our model improves by an average of 1.27% on cell-type annotation and 3.86% on perturbation prediction. Furthermore, our method significantly decreases computational overhead and exhibits few-shot potential. GenoHoption can function as an efficient and expressive bridge, connecting existing single-cell foundation models to gene network graphs.

q-bio.QM