Searcharxiv⌕ Search

arXiv subjects

Zahra Khodagholi

Publications and source records attributed to Zahra Khodagholi.

3 recordsLinked to original sources

One pool, many targets: a conservation layer and what archival data can identify

Pairwise guide--transcript scores do not enforce conservation of a finite guide-loaded RISC pool when they are interpreted independently as occupancies. We formulate a differentiable scalar equilibrium layer: one conservation equation with a unique positive root and exact implicit gradients. It yields a redistribution theorem, a qualified high-resource limit, an analysis of the retrieval approximation, and a conditional rank-invariance result: within one construct at one dose, rankings by fractional occupancy cannot distinguish equilibrium from independent scoring. We therefore audit the two experiments that proposition leaves open, dose and cross-context, on archival off-target data. Corrected thermodynamic affinities associate weakly with measured repression in the direction a working predictor requires, but a paired permutation test and a construct-cluster bootstrap do not establish added predictive value from the coupling: what survives their differing permutation-null baselines is \GapNet{}, a descriptive \GapNetOverSE{} of the equilibrium association's cluster standard error. The dose fits are heterogeneous and frequently violate the model-implied exponent constraint, which is superlinear rather than sublinear, so these data do not identify the competition parameter. A saturable compression of the competitor set holds both accuracy targets on held-out guide families but is not faster at the size measured. The contribution is a reusable conservation operator and the experimental information needed to test it. The code for this study is available at https://github.com/shadi97kh/One-Pool-Many-Targets.

cs.LG↗

Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor

Input encodings can restrict which measured contrasts a predictor can jointly reproduce, even when no single contrast is forced to vanish. We compute the attainable contrast space from an encoder's equivalence classes and a fixed contrast design, without labels, loss, or a fitted model; projecting the recorded contrasts onto that space gives an empirical error floor for any unrestricted decoder on those classes. On a 140-rectangle siRNA interaction panel, a graph neural network's training-only feature mask merges 165 endpoint states into 90 classes and cuts the rank of the 140 interaction contrasts to 72. The resulting floor is 0.009980, which is 14.6% of the fitted model's interaction squared error; the fitted model reaches 0.068335, slightly worse than a control predicting no interaction at all. A minimum of three restored chemistry columns recovers full rank. Refitting without the mask removes the floor entirely, yet interaction MSE improves by only 0.000017 under the reported protocol, and the restored columns remain absent from every training input. On a released RNA-splicing predictor, whose encoding is injective on the measured states, the same computation returns the full design rank of 1,986 and a floor of exactly zero. These results separate what an encoding permits from what a fitted model achieves; they do not identify what limits the remaining error. The rank check needs no fits and bounds what any amount of training under a fixed encoding can recover. The project repository is available at https://github.com/shadi97kh/REPRESENTABLE-BUT-UNLEARNED.

cs.LG↗

Validating Interpretability in siRNA Efficacy Prediction: A Perturbation-Based, Dataset-Aware Protocol

Saliency maps are increasingly used as design guidance in siRNA efficacy prediction, yet attribution methods are rarely validated before motivating sequence edits. We introduce a pre-synthesis gate: a protocol for counterfactual sensitivity faithfulness that tests whether mutating high-saliency positions changes model output more than composition-matched controls. Cross-dataset transfer reveals two failure modes that would otherwise go undetected: faithful-but-wrong (saliency valid, predictions fail) and inverted saliency (top-saliency edits less impactful than random). Strikingly, models trained on mRNA-level assays collapse on a luciferase reporter dataset, demonstrating that protocol shifts can silently invalidate deployment. Across four benchmarks, 19/20 fold instances pass; the single failure shows inverted saliency. A biology-informed regularizer (BioPrior) strengthens saliency faithfulness with modest, dataset-dependent predictive trade-offs. Our results establish saliency validation as essential pre-deployment practice for explanation-guided therapeutic design. Code is available at https://github.com/shadi97kh/BioPrior.

q-bio.GN↗