SearcharxivSearch

arXiv subjects

Alex J. Lee

Publications and source records attributed to Alex J. Lee.

4 recordsLinked to original sources

Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks

The increasing depth of parametric domain knowledge in large language models (LLMs) is fueling their rapid deployment in real-world applications. Understanding model vulnerabilities in high-stakes and knowledge-intensive tasks is essential for quantifying the trustworthiness of model predictions and regulating their use. The recent discovery of named entities as adversarial examples (i.e. adversarial entities) in natural language processing tasks raises questions about their potential impact on the knowledge robustness of pre-trained and finetuned LLMs in high-stakes and specialized domains. We examined the use of type-consistent entity substitution as a template for collecting adversarial entities for billion-parameter LLMs with biomedical knowledge. To this end, we developed an embedding-space attack based on powerscaled distance-weighted sampling to assess the robustness of their biomedical knowledge with a low query budget and controllable coverage. Our method has favorable query efficiency and scaling over alternative approaches based on random sampling and blackbox gradient-guided search, which we demonstrated for adversarial distractor generation in biomedical question answering. Subsequent failure mode analysis uncovered two regimes of adversarial entities on the attack surface with distinct characteristics and we showed that entity substitution attacks can manipulate token-wise Shapley value explanations, which become deceptive in this setting. Our approach complements standard evaluations for high-capacity models and the results highlight the brittleness of domain knowledge in LLMs.

cs.CL

Machine Learning for Uncovering Biological Insights in Spatial Transcriptomics Data

Development and homeostasis in multicellular systems both require exquisite control over spatial molecular pattern formation and maintenance. Advances in spatially-resolved and high-throughput molecular imaging methods such as multiplexed immunofluorescence and spatial transcriptomics (ST) provide exciting new opportunities to augment our fundamental understanding of these processes in health and disease. The large and complex datasets resulting from these techniques, particularly ST, have led to rapid development of innovative machine learning (ML) tools primarily based on deep learning techniques. These ML tools are now increasingly featured in integrated experimental and computational workflows to disentangle signals from noise in complex biological systems. However, it can be difficult to understand and balance the different implicit assumptions and methodologies of a rapidly expanding toolbox of analytical tools in ST. To address this, we summarize major ST analysis goals that ML can help address and current analysis trends. We also describe four major data science concepts and related heuristics that can help guide practitioners in their choices of the right tools for the right biological questions.

q-bio.QM

Accurate Hellmann-Feynman forces from density functional calculations with augmented Gaussian basis sets

The Hellmann-Feynman (HF) theorem provides a way to compute forces directly from the electron density, enabling efficient force calculations for large systems through machine learning (ML) models for the electron density. The main issue holding back the general acceptance of the HF approach for atom-centered basis sets is the well-known Pulay force which, if naively discarded, typically constitutes an error upwards of 10 eV/Ang in forces. In this work, we demonstrate that if a suitably augmented Gaussian basis set is used for density functional calculations, the Pulay force can be suppressed and HF forces can be computed as accurately as analytical forces with state-of-the-art basis sets, allowing geometry optimization and molecular dynamics to be reliably performed with HF forces. Our results pave a clear path forwards for the accurate and efficient simulation of large systems using ML densities and the HF theorem.

physics.chem-ph

Dopant levels in large nanocrystals using stochastic optimally tuned range-separated hybrid density functional theory

We apply a stochastic version of an optimally tuned range-separated hybrid functional to provide insight on the electronic properties of P- and B- doped Si nanocrystals of experimentally relevant sizes. We show that we can use the range-separation parameter for undoped systems to calculate accurate results for dopant activation energies. We apply this strategy for tuning functionals to study doped nanocrystals up to 2.5 nm in diameter at the hybrid functional level. In this confinement regime, the P- and B- dopants have large activation energies and have strongly localized states that lie deep within the energy gaps. Structural relaxation plays a greater role for B-substituted dopants and contributes to the increase in activation energy when the B dopant is near the nanocrystal surface.

cond-mat.mtrl-sci