SearcharxivβŒ• Search

arXiv subjects

Duminda S. Ranasinghe

Publications and source records attributed to Duminda S. Ranasinghe.

2 recordsLinked to original sources

Interpreting GFlowNets for Drug Discovery: What probes can and cannot show

Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting adoption in drug discovery, where chemists need interpretable rationales for proposed structures. We present a control-validated interpretability study of SynFlowNet, a synthesis-aware GFlowNet trained with a drug-likeness (QED) reward. Our framework combines gradient saliency and counterfactual edits, an undercomplete factor analysis, and an overcomplete BatchTopK sparse autoencoder, evaluated with shuffled-label controls, RDKit-descriptor baselines, a scaffold-disjoint split, cross-seed stability, and an architecture-matched untrained network. These controls materially change the interpretation. Physicochemical properties and functional groups are highly decodable from SynFlowNet embeddings, but an untrained network with the same architecture performs essentially as well as the trained policy (drug-likeness within 0.005 and marginally better molecular-size decoding). Thus, this decodability reflects graph architecture and atom featurization rather than representations acquired through policy training. The surviving conclusions are narrower: the overcomplete autoencoder reconstructs embeddings better than a matched undercomplete baseline at comparable sparsity; chemically enriched substructure detectors, including per-halogen and boron features, emerge in individual runs, although only a small subset of dictionary directions is stable across seeds; and zeroing individual features produces property-specific effects on probe decoding. Beyond SynFlowNet, this study provides a reusable control protocol for molecular-model interpretability, separating learned structure from signals supplied by architecture and input representation. High probe scores alone should not be treated as evidence of learned chemistry.

cs.LG↗

Predictive coupled-cluster isomer orderings for some Si${}_n$C${}_m$ ($m, n\le 12$) clusters; A pragmatic comparison between DFT and complete basis limit coupled-cluster benchmarks

The accurate determination of the preferred ${\rm Si}_{12}{\rm C}_{12}$ isomer is important to guide experimental efforts directed towards synthesizing SiC nano-wires and related polymer structures which are anticipated to be highly efficient exciton materials for opto-electronic devices. In order to definitively identify preferred isomeric structures for silicon carbon nano-clusters, highly accurate geometries, energies and harmonic zero point energies have been computed using coupled-cluster theory with systematic extrapolation to the complete basis limit for set of silicon carbon clusters ranging in size from SiC$_3$ to ${\rm Si}_{12}{\rm C}_{12}$. It is found that post-MBPT(2) correlation energy plays a significant role in obtaining converged relative isomer energies, suggesting that predictions using low rung density functional methods will not have adequate accuracy. Utilizing the best composite coupled-cluster energy that is still computationally feasible, entailing a 3-4 SCF and CCSD extrapolation with triple-$ΞΆ$ (T) correlation, the {\it closo} ${\rm Si}_{12}{\rm C}_{12}$ isomer is identified to be the preferred isomer in support of previous calculations [J. Chem. Phys. 2015, 142, 034303]. Additionally we have investigated more pragmatic approaches to obtaining accurate silicon carbide isomer energies, including the use of frozen natural orbital coupled-cluster theory and several rungs of standard and double-hybrid density functional theory. Frozen natural orbitals as a way to compute post MBPT(2) correlation energy is found to be an excellent balance between efficiency and accuracy.

physics.chem-ph↗