SearcharxivSearch

arXiv subjects

Ning Sui

Publications and source records attributed to Ning Sui.

9 recordsLinked to original sources

PlantBGC: Transformer for Plant BGC Discovery via Label-Free Domain Adaptation and Weak Supervision

Plant biosynthetic gene clusters (BGCs) encode specialized-metabolite pathways, yet curated plant BGC labels remain scarce, hindering supervised discovery at genome scale. Existing plant BGC mining tools are largely signature- and rule-driven and do not fully leverage recent advances in contextual representation learning for modeling long-range domain context and controlling false positives under strong domain shift. We seek an AI-assisted workflow that narrows experimental search space by transferring supervision from well-annotated microbial BGCs to plant genomes. We present PlantBGC, representing genomes as ordered Pfam-domain sequences and learning BGC-likeness with an encoder-only Transformer trained on MIBiG microbial BGCs and adapted to plants via label-free masked language modeling. On microbial benchmarks, PlantBGC achieves token-level AUC = 0.988 (10-fold CV) and 0.979 (leave-class-out). On plants, adaptation improves known-BGC recovery on n = 34 curated loci under strict 100% coverage, increasing recovery from 29.4% to 67.6% and indicating more complete boundaries. GO/KEGG-derived weak supervision reduces proxy primary-like ratio by 48.40% (GO) and 45.20% (KEGG), with consistent per-species reductions (paired Wilcoxon p = 1.53e-5). Compared to plantiSMASH, PlantBGC yields more compact loci on matched regions (median length ratio = 0.278; 93.8% of pairs are shorter).

q-bio.GN

Faculty Orientations Shape Adoption of AI in Research and Teaching

Despite the widespread availability of large language models (LLMs) in higher education, instructors vary substantially in their adoption and use of these tools, and the reasons for this variation remain poorly understood. A mixed-methods survey of 90 STEM faculty in the Research Corporation for Science Advancement (RCSA) Cottrell community examined relationships between AI use, attitudes, institutional context, and instructional practice. Exploratory factor analysis identified a coherent construct, \textit{AI pedagogical orientation}, that strongly predicted self-reported AI use across research, teaching, and other professional activities. Qualitative analysis indicated that this construct reflected differing views about the role AI should play in disciplinary thinking, learning, and expertise development, rather than simply positive or negative attitudes toward AI. Institutional initiatives, demographic variables, and information sources showed comparatively weak associations with AI use. The results suggest that existing technology-adoption models may not fully explain adoption in contexts where technologies interact directly with disciplinary reasoning and knowledge production.

physics.ed-ph

DeepDTF: Dual-Branch Transformer Fusion for Multi-Omics Anticancer Drug Response Prediction

Cancer drug response varies widely across tumors due to multi-layer molecular heterogeneity, motivating computational decision support for precision oncology. Despite recent progress in deep CDR models, robust alignment between high-dimensional multi-omics and chemically structured drugs remains challenging due to cross-modal misalignment and limited inductive bias. We present DeepDTF, an end-to-end dual-branch Transformer fusion framework for joint log(IC50) regression and drug sensitivity classification. The cell-line branch uses modality-specific encoders for multi-omics profiles with Transformer blocks to capture long-range dependencies, while the drug branch represents compounds as molecular graphs and encodes them with a GNN-Transformer to integrate local topology with global context. Omics and drug representations are fused by a Transformer-based module that models cross-modal interactions and mitigates feature misalignment. On public pharmacogenomic benchmarks under 5-fold cold-start cell-line evaluation, DeepDTF consistently outperforms strong baselines across omics settings, achieving up to RMSE=1.248, R^2=0.875, and AUC=0.987 with full multi-omics inputs, while reducing classification error (1-ACC) by 9.5%. Beyond accuracy, DeepDTF provides biologically grounded explanations via SHAP-based gene attributions and pathway enrichment with pre-ranked GSEA.

cs.LG

Pre-trained Molecular Language Models with Random Functional Group Masking

Recent advancements in computational chemistry have leveraged the power of trans-former-based language models, such as MoLFormer, pre-trained using a vast amount of simplified molecular-input line-entry system (SMILES) sequences, to understand and predict molecular properties and activities, a critical step in fields like drug discovery and materials science. To further improve performance, researchers have introduced graph neural networks with graph-based molecular representations, such as GEM, incorporating the topology, geometry, 2D or even 3D structures of molecules into pre-training. While most of molecular graphs in existing studies were automatically converted from SMILES sequences, it is to assume that transformer-based language models might be able to implicitly learn structure-aware representations from SMILES sequences. In this paper, we propose \ours{} -- a SMILES-based \underline{\em M}olecular \underline{\em L}anguage \underline{\em M}odel, which randomly masking SMILES subsequences corresponding to specific molecular \underline{\em F}unctional \underline{\em G}roups to incorporate structure information of atoms during the pre-training phase. This technique aims to compel the model to better infer molecular structures and properties, thus enhancing its predictive capabilities. Extensive experimental evaluations across 11 benchmark classification and regression tasks in the chemical domain demonstrate the robustness and superiority of \ours{}. Our findings reveal that \ours{} outperforms existing pre-training models, either based on SMILES or graphs, in 9 out of the 11 downstream tasks, ranking as a close second in the remaining ones.

q-bio.BM

Stellar outbursts and chondrite composition

The temperatures of observed protoplanetary disks are not sufficiently high to produce the accretion rate needed to form stars, nor are they sufficient to explain the volatile depletion patterns in CM, CO, and CV chondrites and terrestrial planets. We revisit the role that stellar outbursts, caused by high accretion episodes, play in resolving these two issues. These outbursts provide the necessary mass to form the star during the disk lifetime, and provide enough heat to vaporize planet-forming materials. We show that these outbursts can reproduce the observed chondrite abundances at distances near one AU. These outbursts would also affect the growth of calcium-aluminum-rich inclusions (CAIs) and the isotopic compositions of carbonaceous and non-carbonaceous chondrites.

astro-ph.EP

Protoformer: Embedding Prototypes for Transformers

Transformers have been widely applied in text classification. Unfortunately, real-world data contain anomalies and noisy labels that cause challenges for state-of-art Transformers. This paper proposes Protoformer, a novel self-learning framework for Transformers that can leverage problematic samples for text classification. Protoformer features a selection mechanism for embedding samples that allows us to efficiently extract and utilize anomalies prototypes and difficult class prototypes. We demonstrated such capabilities on datasets with diverse textual structures (e.g., Twitter, IMDB, ArXiv). We also applied the framework to several models. The results indicate that Protoformer can improve current Transformers in various empirical settings.

cs.CL

On the Gravitational Instabilities of Protoplanetary Disks

The gravitational instabilities are important to the evolution of the disks and the planet formation in the disks. We calculate the evolution of the disks which form from the collapse of the molecular cloud cores. By changing the properties of the cloud cores and the hydrodynamical viscosity parameters, we explore their effects on the properties of the gravitational instabilities. We find that the disk is unstable when the angular velocity of the molecular cloud core is larger than a critical value. The time duration of the instability increases as the angular velocity of the core increases. The increase of the hydrodynamical viscosity parameter hardly affects the stability of the disk, but decreases the time duration of the critical state of the gravitational instability in the disk. The instability of the disks can happen at very early time of evolution of the disk, which is consistent with the observations.

astro-ph.EP

Constraining the optical depth of galaxies and velocity bias with cross-correlation between kinetic Sunyaev-Zeldovich effect and peculiar velocity field

We calculate the cross-correlation function $\langle (ΔT/T)(\mathbf{v}\cdot \mathbf{n}/σ_{v}) \rangle$ between the kinetic Sunyaev-Zeldovich (kSZ) effect and the reconstructed peculiar velocity field using linear perturbation theory, to constrain the optical depth $τ$ and peculiar velocity bias of central galaxies with Planck data. We vary the optical depth $τ$ and the velocity bias function $b_{v}(k)=1+b(k/k_{0})^{n}$, and fit the model to the data, with and without varying the calibration parameter $y_{0}$ that controls the vertical shift of the correlation function. By constructing a likelihood function and constraining $τ$, $b$ and $n$ parameters, we find that the quadratic power-law model of velocity bias $b_{v}(k)=1+b(k/k_{0})^{2}$ provides the best-fit to the data. The best-fit values are $τ=(1.18 \pm 0.24) \times 10^{-4}$, $b=-0.84^{+0.16}_{-0.20}$ and $y_{0}=(12.39^{+3.65}_{-3.66})\times 10^{-9}$ ($68\%$ confidence level). The probability of $b>0$ is only $3.12 \times 10^{-8}$ for the parameter $b$, which clearly suggests a detection of scale-dependent velocity bias. The fitting results indicate that the large-scale ($k \leq 0.1\,h\,{\rm Mpc}^{-1}$) velocity bias is unity, while on small scales the bias tends to become negative. The value of $τ$ is consistent with the stellar mass--halo mass and optical depth relation proposed in the previous literatures, and the negative velocity bias on small scales is consistent with the peak background-split theory. Our method provides a direct tool to study the gaseous and kinematic properties of galaxies.

astro-ph.CO

Statistical computation of Boltzmann entropy and estimation of the optimal probability density function from statistical sample

In this work, we investigate the statistical computation of the Boltzmann entropy of statistical samples. For this purpose, we use both histogram and kernel function to estimate the probability density function of statistical samples. We find that, due to coarse-graining, the entropy is a monotonic increasing function of the bin width for histogram or bandwidth for kernel estimation, which seems to be difficult to select an optimal bin width/bandwidth for computing the entropy. Fortunately, we notice that there exists a minimum of the first derivative of entropy for both histogram and kernel estimation, and this minimum point of the first derivative asymptotically points to the optimal bin width or bandwidth. We have verified these findings by large amounts of numerical experiments. Hence, we suggest that the minimum of the first derivative of entropy be used as a selector for the optimal bin width or bandwidth of density estimation. Moreover, the optimal bandwidth selected by the minimum of the first derivative of entropy is purely data-based, independent of the unknown underlying probability density distribution, which is obviously superior to the existing estimators. Our results are not restricted to one-dimensional, but can also be extended to multivariate cases. It should be emphasized, however, that we do not provide a robust mathematical proof of these findings, and we leave these issues with those who are interested in them.

stat.ME