SearcharxivSearch

arXiv subjects

Aijing Feng

Publications and source records attributed to Aijing Feng.

2 recordsLinked to original sources

Mamba-SSM with LLM Reasoning for Feature Selection: Faithfulness-Aware Biomarker Discovery

Gradient saliency from deep sequence models surfaces candidate biomarkers efficiently, but the resulting gene lists can be contaminated by tissue-composition confounders that degrade downstream classifiers. We study whether LLM chain-of-thought (CoT) reasoning can filter these confounders, and whether reasoning quality is associated with downstream performance. We train a Mamba SSM on TCGA-BRCA RNA-seq and extract the top-50 genes by gradient saliency; DeepSeek-R1 evaluates every candidate with structured CoT to produce a final 17-gene set. On the held-out test split, the raw 50-gene saliency set (no LLM) performs worse than a 5,000-gene variance baseline (AUC 0.832 vs. 0.903), while the LLM-filtered set surpasses it (AUC 0.927), using 294x fewer features. A faithfulness audit (COSMIC CGC, OncoKB, PAM50) shows that 6 of 17 selected genes (35.3%) are validated BRCA biomarkers, while 10 of 16 known BRCA genes present in the input were missed - including FOXA1. This divergence between downstream performance and reasoning faithfulness suggests selective faithfulness in this setting: targeted confounder removal can improve predictive performance without comprehensive recall.

q-bio.QM

SMDP: SARS-CoV-2 Mutation Distribution Profiler for rapid estimation of mutational histories of unusual lineages

SARS-CoV-2 usually evolves at a relatively constant rate over time. Occasionally, however, lineages arise with higher-than-expected numbers of mutations given the date of sampling. Such lineages can arise for a variety of reasons, including selection pressures imposed by evolution during a chronic infection or exposure to mutation-inducing drugs like molnupiravir. We have developed an open-source web-based application (SMDP: SARS-CoV-2 Mutation Distribution Profiler; https://eringill.shinyapps.io/covid_mutation_distributions) that compares a list of user-submitted lineage-defining mutations or a FASTA file containing a single genome (from which lineage-defining mutations are calculated) with established mutation distributions including those observed during (1) the first nine months of the pandemic, (2) during the global transmission of Omicron, (3) during the chronic infection of immunocompromised patients, and (4) during zoonotic spillover from humans to deer. The application calculates the most likely distribution for the user's mutation list and displays log likelihoods for all distributions. In addition, the transition:transversion ratio of the user's list is calculated to determine whether there is evidence of exposure to a mutation-inducing drug such as molnupiravir and indicates whether the list contains mutations in the proofreading domain of nsp14 which could lead to a higher-than-expected mutation rate in the lineage. This tool will be useful for public health and researchers seeking to rapidly infer evolutionary histories of SARS-CoV-2 variants, which can aid risk assessment and public health responses.

q-bio.PE