SearcharxivSearch

arXiv subjects

Pratyush Kumar Shukla

Publications and source records attributed to Pratyush Kumar Shukla.

2 recordsLinked to original sources

CLARA: Clarification of Language Ambiguity through Result Analysis for Natural-Language Cancer Genomics Queries

A natural language interface can be used to make cancer genomics databases easier to use, but even if a question is perfectly fluent, its scientific meaning can be ambiguous. We propose CLARA, a framework that represents a question as a typed scientific query specification, considers a few possible interpretations, executes them, and asks for clarification when the estimates diverge. CLARA was assessed on mutation-prevalence contrasts among eight TCGA PanCancer Atlas cohorts and a 30-gene panel. This benchmark consisted of 330 unique executable contrasts varying in mutation scope, assay denominator, and sample context; 115 contrasts were result-sensitive and 215 were result-stable, per the preregistered definition of relative divergence greater than 0.10 or absolute divergence greater than 5 percentage points. An independently implemented pandas execution engine perfectly replicated all 660 results from the SQLite engine. In a separate 120-question LLM-generated, manually vetted language stress test, CLARA recognized all 60 result-sensitive contrasts and needlessly clarified 13 of 60 stable contrasts (accuracy 89.2%, sensitivity/recall 100%, specificity 78.3%). Standalone machine learning had superior overall accuracy (97.5%) but missed one critical contrast. This demonstrates that downstream execution can distinguish consequential from inconsequential ambiguity and reveal an explicit trade-off between safety and burden.

q-bio.GN

Evaluating Conformal Reliability of Pathway-Level Transcriptomic Signatures Under Cross-Cohort Shift in Sepsis Mortality Prediction

Blood transcriptomic profiling enables prognostic modeling by capturing the host immune response at the molecular level. Yet, the within-cohort evaluation strategies employed by many transcriptomic models inadequately reflect deployment across independent hospitals. Outside deployment scenarios introduce a cohort shift that can substantially degrade predictive performance and reliability of uncertainty estimates. We present a framework for evaluating transcriptomic sepsis mortality prediction under realistic cross-cohort deployment, systematically comparing gene-level, pathway-level and hybrid molecular representations. Four publicly available whole-blood transcriptomic cohorts consisting of 936 patients and 248 mortality events were harmonized into a shared 7,660-gene feature space and evaluated under leave-one-cohort-out validation using logistic regression, random forests, XGBoost and LightGBM. Beyond AUROC and AUPRC, model behavior was evaluated via conformal prediction, calibration analysis, selective prediction and the proposed Pathway Stability Index. Gene-level and hybrid representations were found to generally achieve the strongest discriminative performance, whereas pathway-level representations exhibited greater robustness across model families, more reliable uncertainty behavior under cross-cohort shift and stable molecular signatures enriched for immune and host-defense processes identified through Gene Ontology and KEGG enrichment analyses. These findings demonstrate that molecular representation influences not only predictive discrimination but also calibration, uncertainty reliability, biological coherence and transferability under external validation.

q-bio.GN