SearcharxivSearch

arXiv subjects

Ali Saadat

Publications and source records attributed to Ali Saadat.

7 recordsLinked to original sources

Large Language Models for Variant-Centric Functional Evidence Mining

Functional evidence is essential for clinical interpretation of genomic variants, but identifying relevant studies and translating experimental results into structured evidence remains labor intensive. We developed a benchmark based on ClinGen curated annotations to evaluate two large language models (LLMs), a non reasoning model (gpt-4o-mini) and a reasoning model (o4-mini), on tasks relevant to functional evidence curation: (1) abstract screening to determine whether a study reports functional experiments directly testing specific variants, and (2) full text evidence extraction and classification from matched variant-paper pairs, including interpretation of evidence direction and generation of evidence summaries. Starting from ClinGen variants annotated with functional evidence, we processed curator comments with an LLM to extract PubMed identifiers, evidence labels, and narrative, and retrieved titles, abstracts, and open access PDFs to construct variant-paper pairs. In abstract screening, both models achieved high recall (0.88-0.90) with moderate specificity (0.59-0.65). For full text evidence classification under an explicit variant matching gate, o4-mini achieved 96% accuracy and higher specificity (0.83 vs. 0.37) while maintaining high F1 (0.98 vs. 0.96) compared with gpt-4o-mini. We also used an LLM-as-judge protocol to compare model generated evidence summaries with expert curator comments. Finally, we developed AcmGENTIC, an end to end pipeline that expands variant identifiers, retrieves literature via LitVar2, filters abstracts with LLMs, acquires PDFs, performs multimodal evidence extraction, and generates evidence reports for curator review, with optional agentic parsing of figures and tables. Together, this benchmark and pipeline provide a practical framework for scaling functional evidence curation with human in the loop LLM assistance.

q-bio.GN

From Mutation to Degradation: Predicting Nonsense-Mediated Decay with NMDEP

Nonsense-mediated mRNA decay (NMD) is a critical post-transcriptional surveillance mechanism that degrades transcripts with premature termination codons, safeguarding transcriptome integrity and shaping disease phenotypes. However, accurately predicting NMD efficiency remains challenging, as existing models often rely on simplistic rule-based heuristics or limited feature sets, constraining their accuracy and generalizability. Using paired DNA and RNA data from The Cancer Genome Atlas, we benchmark embedding-only models and demonstrate that they underperform compared to a simple rule-based approach. To address this, we develop NMDEP (NMD Efficiency Predictor), an integrative framework that combines optimized rule-based methods, sequence embeddings, and curated biological features, achieving state-of-the-art predictive performance. Through explainable AI, we identify key NMD determinants, reaffirming established factors such as variant position while uncovering novel contributors like ribosome loading. Applied to over 2.9 million simulated stop-gain variants, NMDEP facilitates large-scale mRNA degradation assessments, advancing variant interpretation and disease research.

q-bio.GN

Proteome-wide prediction of mode of inheritance and molecular mechanism underlying genetic diseases using structural interactomics

Genetic diseases can be classified according to their modes of inheritance and their underlying molecular mechanisms. Autosomal dominant disorders often result from DNA variants that cause loss-of-function, gain-of-function, or dominant-negative effects, while autosomal recessive diseases are primarily linked to loss-of-function variants. In this study, we introduce a graph-of-graphs approach that leverages protein-protein interaction networks and high-resolution protein structures to predict the mode of inheritance of diseases caused by variants in autosomal genes, and to classify dominant-associated proteins based on their functional effect. Our approach integrates graph neural networks, structural interactomics and topological network features to provide proteome-wide predictions, thus offering a scalable method for understanding genetic disease mechanisms.

q-bio.QM

DNA Language Model and Interpretable Graph Neural Network Identify Genes and Pathways Involved in Rare Diseases

Identification of causal genes and pathways is a critical step for understanding the genetic underpinnings of rare diseases. We propose novel approaches to gene prioritization and pathway identification using DNA language model, graph neural networks, and genetic algorithm. Using HyenaDNA, a long-range genomic foundation model, we generated dynamic gene embeddings that reflect changes caused by deleterious variants. These gene embeddings were then utilized to identify candidate genes and pathways. We validated our method on a cohort of rare disease patients with partially known genetic diagnosis, demonstrating the re-identification of known causal genes and pathways and the detection of novel candidates. These findings have implications for the prevention and treatment of rare diseases by enabling targeted identification of new drug targets and therapeutic pathways.

q-bio.QM

Fine-tuning the ESM2 protein language model to understand the functional impact of missense variants

Elucidating the functional effect of missense variants is of crucial importance, yet challenging. To understand the impact of such variants, we fine-tuned the ESM2 protein language model to classify 20 protein features at amino acid resolution. We used the resulting models to: 1) identify protein features that are enriched in either pathogenic or benign missense variants, 2) compare the characteristics of proteins with reference or alternate alleles to understand how missense variants affect protein functionality. We show that our model can be used to reclassify some variants of unknown significance. We also demonstrate the usage of our models for understanding the potential effect of variants on protein features.

q-bio.QM

Constrained Bayesian Optimization Using a Lagrange Multiplier Applied to Power Transistor Design

We propose a novel constrained Bayesian Optimization (BO) algorithm optimizing the design process of Laterally-Diffused Metal-Oxide-Semiconductor (LDMOS) transistors while realizing a target Breakdown Voltage (BV). We convert the constrained BO problem into a conventional BO problem using a Lagrange multiplier. Instead of directly optimizing the traditional Figure-of-Merit (FOM), we set the Lagrangian as the objective function of BO. This adaptive objective function with a changeable Lagrange multiplier can address constrained BO problems which have constraints that require costly evaluations, without the need for additional surrogate models to approximate constraints. Our algorithm enables a device designer to set the target BV in the design space, and obtain a device that satisfies the optimized FOM and the target BV constraint automatically. Utilizing this algorithm, we have also explored the physical limits of the FOM for our devices in 30 - 50 V range within the defined design space.

cs.LG

Transition-Metal Nitride Halide Dielectrics for Transition-Metal Dichalcogenide Transistors

Using first-principles calculations, we investigate six transition-metal nitride halides (TMNHs): HfNBr, HfNCl, TiNBr, TiNCl, ZrNBr, and ZrNCl as potential van der Waals (vdW) dielectrics for transition metal dichalcogenide (TMD) channel transistors. We calculate the exfoliation energies and bulk phonon energies and find that the six TMNHs are exfoliable and thermodynamically stable. We calculate both the optical and static dielectric constants in the in-plane and out-of-plane directions for both monolayer and bulk TMNHs. In monolayers, the out-of-plane static dielectric constant ranges from 5.04 (ZrNCl) to 6.03 (ZrNBr) whereas in-plane dielectric constants range from 13.18 (HfNBr) to 74.52 (TiNCl). We show that the bandgap of TMNHs ranges from 1.53 eV (TiNBr) to 3.36 eV (HfNCl) whereas the affinity ranges from 4.01 eV (HfNBr) to 5.60 eV (TiNCl). Finally, we estimate the dielectric leakage current density of transistors with six TMNH monolayer dielectrics with five monolayer channel TMDs (MoS2, MoSe2, MoTe2, WS2, and WSe2). For p-MOS TMD channel transistors, 19 out of 30 combinations have a smaller leakage current compared to monolayer hexagonal boron nitride (hBN), a well-known vdW dielectric. The smallest monolayer leakage current of 2.14*10-9 A/cm2 is predicted for a p-MOS WS2 transistor with HfNCl as a gate dielectric. HfNBr, HfNCl, ZrNBr, and ZrNCl are also predicted to yield small leakage currents in certain p-MOS TMD transistors.

cond-mat.mtrl-sci