SearcharxivSearch

arXiv subjects

Andrew White

Publications and source records attributed to Andrew White.

17 recordsLinked to original sources

CandorMD: An AI-Assisted Audio Simulation and Feedback System for Training Clinicians for Medical Error Disclosure

Clinicians are expected to disclose harmful medical errors to patients and families in line with ethical, regulatory, and patient care standards, yet these conversations remain challenging because of their emotional complexity and limited training opportunities. Most physicians still learn primarily through lectures and observation, while static video tools-though available-are underused, lack adaptability across specialties, and deliver delayed, generic feedback. These gaps restrict skill development, reduce self-efficacy, and contribute to avoidance of disclosure conversations, ultimately compromising patient care and eroding trust. To address these needs, we designed CandorMD -- an AI-assisted simulation system that provides real-time practice, actionable feedback, and diverse practice environments tailored to individual learning needs. We conducted semi-structured interviews with physicians, risk managers, patient advocates, and communication experts to understand current practices, identify gaps, and collect feedback on CandorMD. Based on these insights, we present findings and design recommendations for the future of AI-supported medical communication training.

cs.HC

Unifying biomedical knowledge in a modern multimodal graph

Biomedical knowledge graphs (KGs) are widely used in the life sciences, yet many are derived from unstructured documents and therefore lack schema-level constraints, whereas graphs assembled from structured resources are difficult to harmonize into a unified representation. We present OptimusKG, a multimodal biomedical labeled property graph (LPG) built from structured and semi-structured resources to preserve factual, type-specific metadata across molecular, anatomical, clinical, and environmental domains. OptimusKG contains 190,939 nodes across 10 entity types, 21,818,752 edges across 27 edge types, and 67,070,490 property instances encoding 109,665,797 values across 145 distinct property keys, derived from 18 ontologies and controlled vocabularies. The graph enforces a top-level schema for nodes and edges and retains granular, type-specific properties, cross-references, and provenance. We assessed the validity of OptimusKG by evaluating whether graph relationships are supported by evidence from the scientific literature using a multimodal agent, PaperQA3. PaperQA3 identified supporting evidence for 70.0% of sampled edges, whereas 83.4% of sampled false edges received no supporting evidence. Edges without literature support were concentrated in associations derived from experimental and functional genomics resources, suggesting that OptimusKG captures biomedical knowledge that may precede synthesis in the scientific literature. OptimusKG is distributed as Apache Parquet files, providing a standardized resource for graph-based machine learning, knowledge-grounded retrieval with large language models, and biomedical discovery use cases such as hypothesis generation.

cs.AI

Graph AI generates neurological hypotheses validated in molecular, organoid, and clinical systems

Neurological diseases are the leading global cause of disability, yet most lack disease-modifying treatments. We present PROTON, a heterogeneous graph transformer that generates testable hypotheses across molecular, organoid, and clinical systems. To evaluate PROTON, we apply it to Parkinson's disease (PD), bipolar disorder (BD), and Alzheimer's disease (AD). In PD, PROTON linked genetic risk loci to genes essential for dopaminergic neuron survival and predicted pesticides toxic to patient-derived neurons, including the insecticide endosulfan, which ranked within the top 1.29% of predictions. In silico screens performed by PROTON reproduced six genome-wide $\alpha$-synuclein experiments, including a split-ubiquitin yeast two-hybrid system (normalized enrichment score [NES] = 2.30, FDR-adjusted $p < 1 \times 10^{-4}$), an ascorbate peroxidase proximity labeling assay (NES = 2.16, FDR $< 1 \times 10^{-4}$), and a high-depth targeted exome sequencing study in 496 synucleinopathy patients (NES = 2.13, FDR $< 1 \times 10^{-4}$). In BD, PROTON predicted calcitriol as a candidate drug that reversed proteomic alterations observed in cortical organoids derived from BD patients. In AD, we evaluated PROTON predictions in health records from $n = 610,524$ patients at Mass General Brigham, confirming that five PROTON-predicted drugs were associated with reduced seven-year dementia risk (minimum hazard ratio = 0.63, 95% CI: 0.53-0.75, $p < 1 \times 10^{-7}$). PROTON generated neurological hypotheses that were evaluated across molecular, organoid, and clinical systems, defining a path for AI-driven discovery in neurological disease.

q-bio.QM

BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology

Large Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research. Existing benchmarks for measuring this potential and guiding future development continue to evolve from pure recall and rote knowledge tasks, towards more practical work such as literature review and experimental planning. Bioinformatics is a domain where fully autonomous AI-driven discovery may be near, but no extensive benchmarks for measuring progress have been introduced to date. We therefore present the Bioinformatics Benchmark (BixBench), a dataset comprising over 50 real-world scenarios of practical biological data analysis with nearly 300 associated open-answer questions designed to measure the ability of LLM-based agents to explore biological datasets, perform long, multi-step analytical trajectories, and interpret the nuanced results of those analyses. We evaluate the performance of two frontier LLMs (GPT-4o and Claude 3.5 Sonnet) using a custom agent framework we open source. We find that even the latest frontier models only achieve 17% accuracy in the open-answer regime, and no better than random in a multiple-choice setting. By exposing the current limitations of frontier models, we hope BixBench can spur the development of agents capable of conducting rigorous bioinformatic analysis and accelerate scientific discovery.

q-bio.QM

Report of the Topical Group on Calorimetry

The 2022 Snowmass Community Summer Study consisted of two years of engagement with the particle physics community to craft a US particle physics strategy for the next decade, culminating in the 2022 Snowmass Book. This Report of the Topical Group on Calorimetry forms the basis for a small section of the Snowmass Book dealing with the community's vision for calorimetry going forward. It is the distillation of ideas put forth in public meetings and numerous white papers already freely available. The Report describes particle flow and dual readout approaches, integration of precision timing, and trends in calorimetric materials, and highlights key research directions for the next decade.

physics.ins-det

SELFIES and the future of molecular string representations

Artificial intelligence (AI) and machine learning (ML) are expanding in popularity for broad applications to challenging tasks in chemistry and materials science. Examples include the prediction of properties, the discovery of new reaction pathways, or the design of new molecules. The machine needs to read and write fluently in a chemical language for each of these tasks. Strings are a common tool to represent molecular graphs, and the most popular molecular string representation, SMILES, has powered cheminformatics since the late 1980s. However, in the context of AI and ML in chemistry, SMILES has several shortcomings -- most pertinently, most combinations of symbols lead to invalid results with no valid chemical interpretation. To overcome this issue, a new language for molecules was introduced in 2020 that guarantees 100\% robustness: SELFIES (SELF-referencIng Embedded Strings). SELFIES has since simplified and enabled numerous new applications in chemistry. In this manuscript, we look to the future and discuss molecular string representations, along with their respective opportunities and challenges. We propose 16 concrete Future Projects for robust molecular representations. These involve the extension toward new chemical domains, exciting questions at the interface of AI and robust languages and interpretability for both humans and machines. We hope that these proposals will inspire several follow-up works exploiting the full potential of molecular string representations for the future of AI in chemistry and materials science.

physics.chem-ph

The International Linear Collider: Report to Snowmass 2021

The International Linear Collider (ILC) is on the table now as a new global energy-frontier accelerator laboratory taking data in the 2030s. The ILC addresses key questions for our current understanding of particle physics. It is based on a proven accelerator technology. Its experiments will challenge the Standard Model of particle physics and will provide a new window to look beyond it. This document brings the story of the ILC up to date, emphasizing its strong physics motivation, its readiness for construction, and the opportunity it presents to the US and the global particle physics community.

physics.acc-ph

Federated Learning of Molecular Properties with Graph Neural Networks in a Heterogeneous Setting

Chemistry research has both high material and computational costs to conduct experiments. Institutions thus consider chemical data to be valuable and there have been few efforts to construct large public datasets for machine learning. Another challenge is that different intuitions are interested in different classes of molecules, creating heterogeneous data that cannot be easily joined by conventional distributed training. In this work, we introduce federated heterogeneous molecular learning to address these challenges. Federated learning allows end-users to build a global model collaboratively while keeping the training data distributed over isolated clients. Due to the lack of related research, we first simulate a heterogeneous federated learning benchmark (FedChem) by jointly performing scaffold splitting and latent Dirichlet allocation on existing datasets for heterogeneously distributed client data. Our results on FedChem show that significant learning challenges arise when working with heterogeneous molecules across clients. We then propose a method to alleviate the problem, namely Federated Learning by Instance reweighTing (FLIT(+)). FLIT(+) can align the local training across heterogeneous clients by improving the performance for uncertain samples. Comprehensive experiments conducted on our new benchmark FedChem validate the advantages of this method over other federated learning schemes. FedChem should enable a new type of collaboration for improving AI in chemistry that mitigates concerns about valuable chemical data.

cs.LG

The International Linear Collider: A Global Project

The International Linear Collider (ILC) is now under consideration as the next global project in particle physics. In this report, we review of all aspects of the ILC program: the physics motivation, the accelerator design, the run plan, the proposed detectors, the experimental measurements on the Higgs boson, the top quark, the couplings of the W and Z bosons, and searches for new particles. We review the important role that polarized beams play in the ILC program. The first stage of the ILC is planned to be a Higgs factory at 250 GeV in the centre of mass. Energy upgrades can naturally be implemented based on the concept of a linear collider. We discuss in detail the ILC program of Higgs boson measurements and the expected precision in the determination of Higgs couplings. We compare the ILC capabilities to those of the HL-LHC and to those of other proposed e+e- Higgs factories. We emphasize throughout that the readiness of the accelerator and the estimates of ILC performance are based on detailed simulations backed by extensive RandD and, for the accelerator technology, operational experience.

hep-ex

The International Linear Collider. A Global Project

A large, world-wide community of physicists is working to realise an exceptional physics program of energy-frontier, electron-positron collisions with the International Linear Collider (ILC). This program will begin with a central focus on high-precision and model-independent measurements of the Higgs boson couplings. This method of searching for new physics beyond the Standard Model is orthogonal to and complements the LHC physics program. The ILC at 250 GeV will also search for direct new physics in exotic Higgs decays and in pair-production of weakly interacting particles. Polarised electron and positron beams add unique opportunities to the physics reach. The ILC can be upgraded to higher energy, enabling precision studies of the top quark and measurement of the top Yukawa coupling and the Higgs self-coupling. The key accelerator technology, superconducting radio-frequency cavities, has matured. Optimised collider and detector designs, and associated physics analyses, were presented in the ILC Technical Design Report, signed by 2400 scientists. There is a strong interest in Japan to host this international effort. A detailed review of the many aspects of the project is nearing a conclusion in Japan. Now the Japanese government is preparing for a decision on the next phase of international negotiations, that could lead to a project start within a few years. The potential timeline of the ILC project includes an initial phase of about 4 years to obtain international agreements, complete engineering design and prepare construction, and form the requisite international collaboration, followed by a construction phase of 9 years.

hep-ex

The limitations of model-based experimental design and parameter estimation in sloppy systems

We explore the relationship among model fidelity, experimental design, and parameter estimation in sloppy models. We show that the approximate nature of mathematical models poses challenges for experimental design in sloppy models. In many models of complex biological processes it is unknown what are the relevant physics that must be included to explain collective behaviors. As a consequence, models are often overly complex, with many practically unidentifiable parameters. Furthermore, which details are relevant/irrelevant vary among potential experiments. By selecting complementary experiments, experimental design may inadvertently make details that were ommitted from the model become relevant. When this occurs, the model will fail to give a good fit to the data. We use a simple hyper-model of model error to quantify a model's inadequacy and apply it to two models of complex biological processes (EGFR signaling and DNA repair) with optimally selected experiments. We find that although parameters may be accurately estimated, the error in the model renders it less predictive than it was in the sloppy regime where model error is small. We introduce the concept of a \emph{sloppy system}--a sequence of models of increasing complexity that become sloppy in the limit of microscopic accuracy. We explore the limits of accurate parameter estimation in sloppy systems and argue that system identification better approached by considering a hierarchy of models of varying detail rather than focusing parameter estimation in a single model.

q-bio.QM

An Easy to Use Repository for Comparing and Improving Machine Learning Algorithm Usage

The results from most machine learning experiments are used for a specific purpose and then discarded. This results in a significant loss of information and requires rerunning experiments to compare learning algorithms. This also requires implementation of another algorithm for comparison, that may not always be correctly implemented. By storing the results from previous experiments, machine learning algorithms can be compared easily and the knowledge gained from them can be used to improve their performance. The purpose of this work is to provide easy access to previous experimental results for learning and comparison. These stored results are comprehensive -- storing the prediction for each test instance as well as the learning algorithm, hyperparameters, and training set that were used. Previous results are particularly important for meta-learning, which, in a broad sense, is the process of learning from previous machine learning results such that the learning process is improved. While other experiment databases do exist, one of our focuses is on easy access to the data. We provide meta-learning data sets that are ready to be downloaded for meta-learning experiments. In addition, queries to the underlying database can be made if specific information is desired. We also differ from previous experiment databases in that our databases is designed at the instance level, where an instance is an example in a data set. We store the predictions of a learning algorithm trained on a specific training set for each instance in the test set. Data set level information can then be obtained by aggregating the results from the instances. The instance level information can be used for many tasks such as determining the diversity of a classifier or algorithmically determining the optimal subset of training instances for a learning algorithm.

stat.ML

Instrumentation for the Energy Frontier

The Instrumentation Frontier was set up as a part of the Snowmass 2013 Community Summer Study to examine the instrumentation R&D needed to support particle physics research over the coming decade. This report summarizes the findings of the Energy Frontier subgroup of the Instrumentation Frontier.

physics.ins-det

Levitating Drop in a Tilted Rotating Tank - Gallery of Fluid Motion Entry V044

A cylindrical acrylic tank with inner diameter D = 4 in. is mounted such that its axis of symmetry is at some angle measured from the vertical plane. The mixing tank is identical to that described in [1] The tank is filled with 200 mL of 1000 cSt silicone oil and a 5 mL drop of de-ionized water is placed in the oil volume. The water drop is allowed to come to rest and then a motor rotates the tank about its axis of symmetry at a fixed frequency = 0.3 Hz. Therefore the Reynolds number is fixed at about Re ~ 5 yielding laminar flow conditions. A CCD camera (PixeLink) is used to capture video of each experiment.

physics.flu-dyn

APS DFD 2011 video submission V045

Inhomogeneous uid mixing in a tilted-rotating cylindrical tank (radius a = 3:5 cm) is shown at Re(17-40) and low capillary numbers. A water and surfactant solu- tion (1% by mass sodium oleate) is dispersed in soybean oil (95% by volume), through varying the rotation rate, and angle of inclination, the rate of mixing is observed. A planar laser is directed down the tank axis to highlight a cross-sectional area of the fluid volume and as the water droplets begin to break up to sizes on the order of the beam width and less, more light is refracted and the mixture is illuminated. Initially, the water breaks up into large droplets that exhibit approximate solid-body rotation about the bottom of the tank. When the total combined volume is below the critical volume of the tank Vcrit = a^3 tan vortex transport of the water occurs more rapidly, breaking up the water into continually smaller droplets in a process that resembles periodic shearing. When the fluid volume is above critical the water will break up and rotate about the bottom of the tank and vortex-induced mixing is much more reticent, if occurring at all. It is noted that shallower angles with respect to the horizontal produce faster mixing while allowing a greater volume of fluid to be mixed at the sub-critical volume given a constant tank size.

physics.flu-dyn

Life history and mating systems select for male biased parasitism mediated through natural selection and ecological feedbacks

Males are often the "sicker" sex with male biased parasitism found in a taxonomically diverse range of species. There is considerable interest in the processes that could underlie the evolution of sex-biased parasitism. Mating system differences along with differences in lifespan may play a key role. We examine whether these factors are likely to lead to male-biased parasitism through natural selection taking into account the critical role that ecological feedbacks play in the evolution of defence. We use a host-parasite model with two-sexes and the techniques of adaptive dynamics to investigate how mating system and sexual differences in competitive ability and longevity can select for a bias in the rates of parasitism. Male-biased parasitism is selected for when males have a shorter average lifespan or when males are subject to greater competition for resources. Male-biased parasitism evolves as a consequence of sexual differences in life history that produce a greater proportion of susceptible females than males and therefore reduce the cost of avoiding parasitism in males. Different mating systems such as monogamy, polygamy or polyandry did not produce a bias in parasitism through these ecological feedbacks but may accentuate an existing bias.

q-bio.PE