SearcharxivSearch

arXiv subjects

Philip Mansfield

Publications and source records attributed to Philip Mansfield.

At least 19 recordsLinked to original sources

Dwarf Galaxies at Cosmic Noon: New JWST Constraints on Satellite Models and Subhalo Tidal Evolution

The advent of JWST has revolutionized the study of faint satellite galaxies at $z \gtrsim 1$, enabling statistical constraints on galaxy evolution and the galaxy$-$halo connection in a previously unexplored mass and redshift regime. We compare satellite abundances at $1 < z < 3.5$ from recent JWST observations with predictions from cosmological dark matter-only zoom-in simulations. We identify and quantify several sources of biases that can impact theoretical satellite counts, finding that assumptions about subhalo tidal evolution introduce the largest uncertainty in predictions for the satellite mass function. Using a flexible galaxy disruption model, we explore a range of disruption scenarios, spanning hydrodynamically motivated and idealized prescriptions, to bracket plausible physical outcomes. We show that varying galaxy durability can change the predicted satellite mass functions by a factor of $\sim3.5$. The JWST data and our fiducial model are consistent within $1-2\sigma$ across the full redshift ($1 < z < 3.5$) and stellar mass ($M_\star> 10^7~\mathrm{M}_\odot$) range probed. We find evidence that subhalos are at least as long-lived as predicted by hydrodynamic simulations. Our framework will enable robust constraints on the tidal evolution of subhalos with future observations. This work presents the first direct comparison between cosmological models and observations of the high-redshift satellite population in this low-mass regime. These results showcase JWST's emerging power to test structure formation in the first half of the Universe in a new domain and to constrain the physical processes driving the evolution of low-mass galaxies across cosmic time.

astro-ph.GA

Novel Challenges in Tracking Self-Interacting Dark Matter Subhalos

Cosmological N-body simulations are among the primary tools for studying structure formation in the Universe. Analyses of these simulations critically depend on accurately identifying and tracking dark matter subhalos over time. In recent years, several new algorithms have been developed to improve the accuracy and consistency of subhalo tracking in cold dark matter simulations. These algorithms should be revisited in the context of new physics beyond gravity, which can modify the evolution and final properties of subhalo populations. In this work, we apply the particle-tracking-based subhalo finder Symfind to velocity-dependent self-interacting dark matter simulations with large cross section amplitudes to assess the performance of particle-tracking methods beyond the CDM paradigm. We find that the core-particle-tracking technique, which is key to the success of these algorithms in CDM, does not always yield accurate results in SIDM. The interplay between dark matter self-interactions and tidal stripping can cause the diffusion of core particles to larger radii, leading particle-tracking-based algorithms to prematurely lose track of SIDM subhalos. For massive core-expansion subhalos and core-collapse subhalos that experience close or repeated pericentric passages, a significant fraction of core particles can be lost, and particle-tracking-based finders such as Symfind offer no clear advantage over traditional methods that rely on identifying phase-space overdensities. On the other hand, for subhalos with large pericentric distances or fewer, more distant passages, Symfind tends to outperform. These differences depend sensitively on the cross section amplitude and turnover velocity of the underlying SIDM model. We therefore recommend a hybrid approach that leverages the strengths of both techniques to produce complete and robust catalogs of core-expansion and core-collapse SIDM subhalos.

astro-ph.CO

Exploring Large Language Models for Specialist-level Oncology Care

Large language models (LLMs) have shown remarkable progress in encoding clinical knowledge and responding to complex medical queries with appropriate clinical reasoning. However, their applicability in subspecialist or complex medical settings remains underexplored. In this work, we probe the performance of AMIE, a research conversational diagnostic AI system, in the subspecialist domain of breast oncology care without specific fine-tuning to this challenging domain. To perform this evaluation, we curated a set of 50 synthetic breast cancer vignettes representing a range of treatment-naive and treatment-refractory cases and mirroring the key information available to a multidisciplinary tumor board for decision-making (openly released with this work). We developed a detailed clinical rubric for evaluating management plans, including axes such as the quality of case summarization, safety of the proposed care plan, and recommendations for chemotherapy, radiotherapy, surgery and hormonal therapy. To improve performance, we enhanced AMIE with the inference-time ability to perform web search retrieval to gather relevant and up-to-date clinical knowledge and refine its responses with a multi-stage self-critique pipeline. We compare response quality of AMIE with internal medicine trainees, oncology fellows, and general oncology attendings under both automated and specialist clinician evaluations. In our evaluations, AMIE outperformed trainees and fellows demonstrating the potential of the system in this challenging and important domain. We further demonstrate through qualitative examples, how systems such as AMIE might facilitate conversational interactions to assist clinicians in their decision making. However, AMIE's performance was overall inferior to attending oncologists suggesting that further research is needed prior to consideration of prospective uses.

cs.HC

Towards Democratization of Subspeciality Medical Expertise

The scarcity of subspecialist medical expertise, particularly in rare, complex and life-threatening diseases, poses a significant challenge for healthcare delivery. This issue is particularly acute in cardiology where timely, accurate management determines outcomes. We explored the potential of AMIE (Articulate Medical Intelligence Explorer), a large language model (LLM)-based experimental AI system optimized for diagnostic dialogue, to potentially augment and support clinical decision-making in this challenging context. We curated a real-world dataset of 204 complex cases from a subspecialist cardiology practice, including results for electrocardiograms, echocardiograms, cardiac MRI, genetic tests, and cardiopulmonary stress tests. We developed a ten-domain evaluation rubric used by subspecialists to evaluate the quality of diagnosis and clinical management plans produced by general cardiologists or AMIE, the latter enhanced with web-search and self-critique capabilities. AMIE was rated superior to general cardiologists for 5 of the 10 domains (with preference ranging from 9% to 20%), and equivalent for the rest. Access to AMIE's response improved cardiologists' overall response quality in 63.7% of cases while lowering quality in just 3.4%. Cardiologists' responses with access to AMIE were superior to cardiologist responses without access to AMIE for all 10 domains. Qualitative examinations suggest AMIE and general cardiologist could complement each other, with AMIE thorough and sensitive, while general cardiologist concise and specific. Overall, our results suggest that specialized medical LLMs have the potential to augment general cardiologists' capabilities by bridging gaps in subspecialty expertise, though further research and validation are essential for wide clinical utility.

cs.HC

EDEN: Exploring Disks Embedded in N-body simulations of Milky-Way-mass halos from Symphony

We investigate the impact of galactic disks on the tidal stripping of cold dark matter subhalos within Milky Way (MW)-mass halos ($M_{\rm vir}\sim 10^{12}\mathrm{M_{\odot}}$) using a new simulation suite, EDEN. By re-simulating 45 MW-mass zoom-in halos from the N-body Symphony compilation with embedded disk potentials, which evolve according to star formation histories predicted by the UniverseMachine model, we self-consistently tie disk growth to halo accretion rate and significantly expand the range of disk masses and formation histories studied. We use the particle-tracking-based subhalo finder Symfind to enhance the robustness of subhalo tracking. We find that disks near the median disk-to-halo mass ratio of our sample ($M_{\ast, \rm Disk}/M_{\rm vir, host} = 2\%$) reduce subhalo peak mass functions within 100 kpc by about $10\%$ for peak masses above $ 10^8\mathrm{M_{\odot}}$. Heavier, MW/M31-like disks ($M_{\ast, \rm Disk}/M_{\rm vir, host} \gtrsim 5\%$) lead to a reduction of more than $40\%$. Subhalo abundance suppression is more pronounced near halo centers, with particularly enhanced stripping for subhalos accreted over 8 Gyr ago on orbits with pericenters < 100 kpc. Suppression is further amplified when disk mass is increased within fixed halo and disk assembly histories. In all cases, the suppression we measure should be interpreted as stripping below the mass resolution limit rather than complete subhalo disruption. This study reshapes our understanding of the MW's impact on its satellites, suggesting it strips subhalos more efficiently than typical MW-mass galaxies due to its larger disk-to-halo mass ratio and earlier disk formation.

astro-ph.GA

Capabilities of Gemini Models in Medicine

Excellence in a wide variety of medical applications poses considerable challenges for AI, requiring advanced reasoning, access to up-to-date medical knowledge and understanding of complex multimodal data. Gemini models, with strong general capabilities in multimodal and long-context reasoning, offer exciting possibilities in medicine. Building on these core strengths of Gemini, we introduce Med-Gemini, a family of highly capable multimodal models that are specialized in medicine with the ability to seamlessly use web search, and that can be efficiently tailored to novel modalities using custom encoders. We evaluate Med-Gemini on 14 medical benchmarks, establishing new state-of-the-art (SoTA) performance on 10 of them, and surpass the GPT-4 model family on every benchmark where a direct comparison is viable, often by a wide margin. On the popular MedQA (USMLE) benchmark, our best-performing Med-Gemini model achieves SoTA performance of 91.1% accuracy, using a novel uncertainty-guided search strategy. On 7 multimodal benchmarks including NEJM Image Challenges and MMMU (health & medicine), Med-Gemini improves over GPT-4V by an average relative margin of 44.5%. We demonstrate the effectiveness of Med-Gemini's long-context capabilities through SoTA performance on a needle-in-a-haystack retrieval task from long de-identified health records and medical video question answering, surpassing prior bespoke methods using only in-context learning. Finally, Med-Gemini's performance suggests real-world utility by surpassing human experts on tasks such as medical text summarization, alongside demonstrations of promising potential for multimodal medical dialogue, medical research and education. Taken together, our results offer compelling evidence for Med-Gemini's potential, although further rigorous evaluation will be crucial before real-world deployment in this safety-critical domain.

cs.AI

A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models

Large language models (LLMs) hold promise to serve complex health information needs but also have the potential to introduce harm and exacerbate health disparities. Reliably evaluating equity-related model failures is a critical step toward developing systems that promote health equity. We present resources and methodologies for surfacing biases with potential to precipitate equity-related harms in long-form, LLM-generated answers to medical questions and conduct a large-scale empirical case study with the Med-PaLM 2 LLM. Our contributions include a multifactorial framework for human assessment of LLM-generated answers for biases, and EquityMedQA, a collection of seven datasets enriched for adversarial queries. Both our human assessment framework and dataset design process are grounded in an iterative participatory approach and review of Med-PaLM 2 answers. Through our empirical study, we find that our approach surfaces biases that may be missed via narrower evaluation approaches. Our experience underscores the importance of using diverse assessment methodologies and involving raters of varying backgrounds and expertise. While our approach is not sufficient to holistically assess whether the deployment of an AI system promotes equitable health outcomes, we hope that it can be leveraged and built upon towards a shared goal of LLMs that promote accessible and equitable healthcare.

cs.CY

Augmentations vs Algorithms: What Works in Self-Supervised Learning

We study the relative effects of data augmentations, pretraining algorithms, and model architectures in Self-Supervised Learning (SSL). While the recent literature in this space leaves the impression that the pretraining algorithm is of critical importance to performance, understanding its effect is complicated by the difficulty in making objective and direct comparisons between methods. We propose a new framework which unifies many seemingly disparate SSL methods into a single shared template. Using this framework, we identify aspects in which methods differ and observe that in addition to changing the pretraining algorithm, many works also use new data augmentations or more powerful model architectures. We compare several popular SSL methods using our framework and find that many algorithmic additions, such as prediction networks or new losses, have a minor impact on downstream task performance (often less than $1\%$), while enhanced augmentation techniques offer more significant performance improvements ($2-4\%$). Our findings challenge the premise that SSL is being driven primarily by algorithmic improvements, and suggest instead a bitter lesson for SSL: that augmentation diversity and data / model scale are more critical contributors to recent advances in self-supervised learning.

cs.LG

Disentangling the Effects of Data Augmentation and Format Transform in Self-Supervised Learning of Image Representations

Self-Supervised Learning (SSL) enables training performant models using limited labeled data. One of the pillars underlying vision SSL is the use of data augmentations/perturbations of the input which do not significantly alter its semantic content. For audio and other temporal signals, augmentations are commonly used alongside format transforms such as Fourier transforms or wavelet transforms. Unlike augmentations, format transforms do not change the information contained in the data; rather, they express the same information in different coordinates. In this paper, we study the effects of format transforms and augmentations both separately and together on vision SSL. We define augmentations in frequency space called Fourier Domain Augmentations (FDA) and show that training SSL models on a combination of these and image augmentations can improve the downstream classification accuracy by up to 1.3% on ImageNet-1K. We also show improvements against SSL baselines in few-shot and transfer learning setups using FDA. Surprisingly, we also observe that format transforms can improve the quality of learned representations even without augmentations; however, the combination of the two techniques yields better quality.

cs.CV

Merger Response of Halo Anisotropy Properties

Anisotropy properties -- halo spin, shape, position offset, velocity offset, and orientation -- are an important family of dark matter halo properties that indicate the level of directional variation of the internal structures of haloes. These properties reflect the dynamical state of haloes, which in turn depends on the mass assembly history. In this work, we study the evolution of anisotropy properties in response to merger activity using the IllustrisTNG simulations. We find that the response trajectories of the anisotropy properties significantly deviate from secular evolution. These trajectories have the same qualitative features and timescales across a wide range of merger and host properties. We propose explanations for the behaviour of these properties and connect their evolution to the relevant stages of merger dynamics. We measure the relevant dynamical timescales. We also explore the dependence of the strength of the response on time of merger, merger ratio, and mass of the main halo. These results provide insight into the physics of halo mergers and their effects on the statistical behaviour of halo properties. This study paves the way towards a physical understanding of scaling relations, particularly to how systematics in their scatter are connected to the mass assembly histories of haloes.

astro-ph.CO

Symfind: Addressing the Fragility of Subhalo Finders and Revealing the Durability of Subhalos

A major question in $\Lambda$CDM is what this theory actually predicts for the properties of subhalo populations. Subhalos are difficult to simulate and to find within simulations, and this propagates into uncertainty in theoretical predictions for satellite galaxies. We present Symfind, a new particle-tracking-based subhalo finder, and demonstrate that it can track subhalos to orders-of-magnitude lower masses than commonly used halo-finding tools, with a focus on Rockstar and consistent-trees. These longer survival mean that at a fixed peak subhalo mass, we find $\approx 15\%{-}40\%$ more subhalos within the virial radius, $R_\textrm{vir}$, and $\approx 35\%-120\%$ more subhalos within $R_\textrm{vir}/4$ in the Symphony dark-matter-only simulation suite. More subhalos are found as resolution is increased. We perform extensive numerical testing. In agreement with idealized simulations, we show that the $v_{\rm max}$ of subhalos is only resolved at high resolutions ($n_\textrm{peak}\gtrsim3\times 10^4$), but that mass loss itself can be resolved at much more modest particle counts ($n_\textrm{peak}\gtrsim4\times 10^3$). We show that Rockstar converges to false solutions for the mass function, radial distribution, and disruption masses of subhalos. We argue that our new method can trace resolved subhalos until the point of typical galaxy disruption without invoking ``orphan'' modeling. We outline a concrete set of steps for determining whether other subhalo finders meet the same criteria. We publicly release Symfind catalogs and particle data for the Symphony simulation suite at \url{web.stanford.edu/group/gfc/gfcsims/}.

astro-ph.CO

Towards Generalist Biomedical AI

Medicine is inherently multimodal, with rich data modalities spanning text, imaging, genomics, and more. Generalist biomedical artificial intelligence (AI) systems that flexibly encode, integrate, and interpret this data at scale can potentially enable impactful applications ranging from scientific discovery to care delivery. To enable the development of these models, we first curate MultiMedBench, a new multimodal biomedical benchmark. MultiMedBench encompasses 14 diverse tasks such as medical question answering, mammography and dermatology image interpretation, radiology report generation and summarization, and genomic variant calling. We then introduce Med-PaLM Multimodal (Med-PaLM M), our proof of concept for a generalist biomedical AI system. Med-PaLM M is a large multimodal generative model that flexibly encodes and interprets biomedical data including clinical language, imaging, and genomics with the same set of model weights. Med-PaLM M reaches performance competitive with or exceeding the state of the art on all MultiMedBench tasks, often surpassing specialist models by a wide margin. We also report examples of zero-shot generalization to novel medical concepts and tasks, positive transfer learning across tasks, and emergent zero-shot medical reasoning. To further probe the capabilities and limitations of Med-PaLM M, we conduct a radiologist evaluation of model-generated (and human) chest X-ray reports and observe encouraging performance across model scales. In a side-by-side ranking on 246 retrospective chest X-rays, clinicians express a pairwise preference for Med-PaLM M reports over those produced by radiologists in up to 40.50% of cases, suggesting potential clinical utility. While considerable work is needed to validate these models in real-world use cases, our results represent a milestone towards the development of generalist biomedical AI systems.

cs.CL

MultiCAM: A multivariable framework for connecting the mass accretion history of haloes with their properties

Models that connect galaxy and halo properties often summarize a halo's mass accretion history (MAH) with a single value, and use this value as the basis for predictions. However, a single-value summary fails to capture the complexity of MAHs and information can be lost in the process. We present MultiCAM, a generalization of traditional abundance matching frameworks, which can simultaneously connect the full MAH of a halo with multiple halo and/or galaxy properties. As a first case study, we apply MultiCAM to the problem of connecting dark matter halo properties to their MAHs in the context of a dark matter-only simulation. While some halo properties, such as concentration, are more strongly correlated to the early-time mass growth of a halo, others, like the virial ratio, have stronger correlations with late-time mass growth. This highlights the necessity of considering the impact of the entire MAH on halo properties. For most of the halo properties we consider, we find that MultiCAM models that use the full MAH achieve higher accuracy than conditional abundance matching models which use a single epoch. We also demonstrate an extension of MultiCAM that captures the covariance between predicted halo properties. This extension provides a baseline model for applications where the covariance between predicted properties is important.

astro-ph.CO

Towards Expert-Level Medical Question Answering with Large Language Models

Recent artificial intelligence (AI) systems have reached milestones in "grand challenges" ranging from Go to protein-folding. The capability to retrieve medical knowledge, reason over it, and answer medical questions comparably to physicians has long been viewed as one such grand challenge. Large language models (LLMs) have catalyzed significant progress in medical question answering; Med-PaLM was the first model to exceed a "passing" score in US Medical Licensing Examination (USMLE) style questions with a score of 67.2% on the MedQA dataset. However, this and other prior work suggested significant room for improvement, especially when models' answers were compared to clinicians' answers. Here we present Med-PaLM 2, which bridges these gaps by leveraging a combination of base LLM improvements (PaLM 2), medical domain finetuning, and prompting strategies including a novel ensemble refinement approach. Med-PaLM 2 scored up to 86.5% on the MedQA dataset, improving upon Med-PaLM by over 19% and setting a new state-of-the-art. We also observed performance approaching or exceeding state-of-the-art across MedMCQA, PubMedQA, and MMLU clinical topics datasets. We performed detailed human evaluations on long-form questions along multiple axes relevant to clinical applications. In pairwise comparative ranking of 1066 consumer medical questions, physicians preferred Med-PaLM 2 answers to those produced by physicians on eight of nine axes pertaining to clinical utility (p < 0.001). We also observed significant improvements compared to Med-PaLM on every evaluation axis (p < 0.001) on newly introduced datasets of 240 long-form "adversarial" questions to probe LLM limitations. While further studies are necessary to validate the efficacy of these models in real-world settings, these results highlight rapid progress towards physician-level performance in medical question answering.

cs.CL

Haunted haloes: tracking the ghosts of subhaloes lost by halo finders

Dark matter subhaloes are key for the predictions of simulations of structure formation, but their existence frequently ends prematurely due to two technical issues, namely numerical disruption in N-body simulations and halo finders failing to identify them. Here we focus on the second issue, using the phase-space friends-of-friends halo finder ROCKSTAR as a benchmark (though we expect our results to translate to comparable codes). We confirm that the most prominent cause for losing track of subhaloes is tidal distortion rather than a low number of particles. As a solution, we present a flexible post-processing algorithm that tracks all subhalo particles over time, computes subhalo positions and masses based on those particles, and progressively removes stripped matter. If a subhalo is lost by the halo finder, this algorithm keeps tracking its so-called ghost until it has almost no particles left or has truly merged with its host. We apply this technique to a large suite of N-body simulations and restore lost subhaloes to the halo catalogues, which has a dramatic effect on key summary statistics of large-scale structure. Specifically, the subhalo mass function increases by about 50% and the halo correlation function increases by a factor of two at small scales. While these quantitative results are somewhat specific to our algorithm, they demonstrate that particle tracking is a promising way to reliably follow haloes and reduce the need for orphan models. Our algorithm and augmented halo catalogues are publicly available.

astro-ph.CO

Symphony: Cosmological Zoom-in Simulation Suites over Four Decades of Host Halo Mass

We present Symphony, a compilation of $262$ cosmological, cold-dark-matter-only zoom-in simulations spanning four decades of host halo mass, from $10^{11}$$-$$10^{15}~M_{\mathrm{\odot}}$. This compilation includes three existing simulation suites at the cluster and Milky Way$-$mass scales, and two new suites: $39$ Large Magellanic Cloud-mass ($10^{11}~M_{\mathrm{\odot}}$) and $49$ strong-lens-analog ($10^{13}~M_{\mathrm{\odot}}$) group-mass hosts. Across the entire host halo mass range, the highest-resolution regions in these simulations are resolved with a dark matter particle mass of $\approx 3\times 10^{-7}$ times the host virial mass and a Plummer-equivalent gravitational softening length of $\approx 9\times 10^{-4}$ times the host virial radius, on average. We measure correlations between subhalo abundance and host concentration, formation time, and maximum subhalo mass, all of which peak at the Milky Way host halo mass scale. Subhalo abundances are $\approx 50\%$ higher in clusters than in lower-mass hosts at fixed sub-to-host halo mass ratios. Subhalo radial distributions are approximately self-similar as a function of host mass and are less concentrated than hosts' underlying dark matter distributions. We compare our results to the semianalytic model $\mathrm{\texttt{Galacticus}}$, which predicts subhalo mass functions with a higher normalization at the low-mass end and radial distributions that are slightly more concentrated than Symphony. We use $\mathrm{\texttt{UniverseMachine}}$ to model halo and subhalo star formation histories in Symphony, and we demonstrate that these predictions resolve the formation histories of the halos that host nearly all currently observable satellite galaxies in the universe. To promote open use of Symphony, data products are publicly available at http://web.stanford.edu/group/gfc/symphony.

astro-ph.CO

Large Language Models Encode Clinical Knowledge

Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but the quality bar for medical and clinical applications is high. Today, attempts to assess models' clinical knowledge typically rely on automated evaluations on limited benchmarks. There is no standard to evaluate model predictions and reasoning across a breadth of tasks. To address this, we present MultiMedQA, a benchmark combining six existing open question answering datasets spanning professional medical exams, research, and consumer queries; and HealthSearchQA, a new free-response dataset of medical questions searched online. We propose a framework for human evaluation of model answers along multiple axes including factuality, precision, possible harm, and bias. In addition, we evaluate PaLM (a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM, on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA, MedMCQA, PubMedQA, MMLU clinical topics), including 67.6% accuracy on MedQA (US Medical License Exam questions), surpassing prior state-of-the-art by over 17%. However, human evaluation reveals key gaps in Flan-PaLM responses. To resolve this we introduce instruction prompt tuning, a parameter-efficient approach for aligning LLMs to new domains using a few exemplars. The resulting model, Med-PaLM, performs encouragingly, but remains inferior to clinicians. We show that comprehension, recall of knowledge, and medical reasoning improve with model scale and instruction prompt tuning, suggesting the potential utility of LLMs in medicine. Our human evaluations reveal important limitations of today's models, reinforcing the importance of both evaluation frameworks and method development in creating safe, helpful LLM models for clinical applications.

cs.CL

Approximating Density Probability Distribution Functions Across Cosmologies

Using a suite of self-similar cosmological simulations, we measure the probability distribution functions (PDFs) of real-space density, redshift-space density, and their geometric mean. We find that the real-space density PDF is well-described by a function of two parameters: $n_s$, the spectral slope, and $σ_L$, the linear rms density fluctuation. For redshift-space density and the geometric mean of real- and redshift-space densities, we introduce a third parameter, $s_L={\sqrt{\langle(dv^L_{\rm pec}/dr)^2\rangle}}/{H}$. We find that density PDFs for the LCDM cosmology is also well-parameterized by these three parameters. As a result, we are able to use a suite of self-similar cosmological simulations to approximate density PDFs for a range of cosmologies. We make the density PDFs publicly available and provide an analytical fitting formula for them.

astro-ph.CO