SearcharxivSearch

arXiv subjects

Maja Jablonska

Publications and source records attributed to Maja Jablonska.

6 recordsLinked to original sources

Milky Way Mapper decoded abundances -- II: From patterns to paths

The element abundances of Milky Way disc stars encode entangled imprints of multiple enrichment processes, making it difficult to uncover the underlying chemical evolution. Here we re-project 16 stellar abundances for 199,290 red giant stars ([Fe/H]$ > -1$) into a set of (4) shared enrichment patterns, providing a generative framework for learning the organising structure of the Milky Way disc. The relative contributions of these patterns vary systematically across the disc, revealing a low-dimensional enrichment basis that responds coherently to global drivers of disc evolution. By grouping stars according to their pattern contributions, we identify coherent enrichment pathways that exhibit strong chemo-spatial correlations and are stratified in both age and height above the plane, linking radial growth to vertical disc structure. Stars occupying similar positions along these enrichment pathways also show coherent vertical deviations across radius, indicating that the low-dimensional chemical structure captures the disc's response to dynamical perturbations. We identify a transition in enrichment behaviour at approximately 6 Gyr, marking the onset of a more chemically mixed regime with increasing contributions from delayed sources. Within this connected system, the observed $\alpha$-bimodality arises within a shared, low-dimensional abundance structure, with stars populating continuous sequences of changing enrichment fractions that are tightly coupled to spatial, temporal, and orbital coordinates across the Milky Way disc.

astro-ph.GA

Milky Way Mapper decoded abundances -- I. Shared disc enrichment patterns

Elemental abundances in the Milky Way disc trace its star-formation and enrichment history, but predicting these abundances from theory is limited by uncertain nucleosynthetic yields and poorly constrained chemical evolution models. Large surveys provide many abundances that enable multi-dimensional insight. However, having so much data available complicates joint visualisation and physical interpretation. Here, we examine the element abundances of 70,057 red giant stars from the Milky Way Mapper survey ([Fe/H] $> -1$), using 16 elements (O,~Mg,~Al,~Si,~S,~K,~Ca,~Ti,~V, ~Cr, Mn,~Fe,~Co,~Ni,~Ce,~Nd). To tackle the challenges of joint-interpretation of these elements, we build a generative data-driven model, expressing each star's abundance vector as a linear combination of a few ($4$) latent nucleosynthetic patterns. These patterns are shared among the population but vary in fraction between stars. The model accurately generates the measured abundances, with $\chi^2 < 3$ (5) for $\sim$ 80\% (95\%) of stars. Model failures, where stars' abundances are not generated by the latent basis reveal accreted material and the role of multiple channels of metal-poor disk enrichment. We associate the recovered patterns, which represent high-precision ($\sigma_P \sim 3$\%) nucleosynthetic channels, with specific enrichment sources; (early and late) core-collapse supernovae, supernovae Type Ia, and asymptotic giant branch stars. We subsequently explore how the dominance of enrichment channels varies across age, metallicity and spatial extent of the disk, and show that enrichment patterns tightly couple to orbital properties. Mean pattern fractions vary smoothly with enrichment, and change rapidly across the valley between the high- and low-$\alpha$ sequences. Our results provide a framework for improving our understanding of Galactic evolution in the Milky Way.

astro-ph.GA

A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models

Hypothesis generation is a fundamental step in scientific discovery, yet it is increasingly challenged by information overload and disciplinary fragmentation. Recent advances in Large Language Models (LLMs) have sparked growing interest in their potential to enhance and automate this process. This paper presents a comprehensive survey of hypothesis generation with LLMs by (i) reviewing existing methods, from simple prompting techniques to more complex frameworks, and proposing a taxonomy that categorizes these approaches; (ii) analyzing techniques for improving hypothesis quality, such as novelty boosting and structured reasoning; (iii) providing an overview of evaluation strategies; and (iv) discussing key challenges and future directions, including multimodal integration and human-AI collaboration. Our survey aims to serve as a reference for researchers exploring LLMs for hypothesis generation.

cs.CL

The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TB of Astronomical Scientific Data

We present the MULTIMODAL UNIVERSE, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, the MULTIMODAL UNIVERSE contains hundreds of millions of astronomical observations, constituting 100\,TB of multi-channel and hyper-spectral images, spectra, multivariate time series, as well as a wide variety of associated scientific measurements and "metadata". In addition, we include a range of benchmark tasks representative of standard practices for machine learning methods in astrophysics. This massive dataset will enable the development of large multi-modal models specifically targeted towards scientific applications. All codes used to compile the MULTIMODAL UNIVERSE and a description of how to access the data is available at https://github.com/MultimodalUniverse/MultimodalUniverse

astro-ph.IM

pathfinder: A Semantic Framework for Literature Review and Knowledge Discovery in Astronomy

The exponential growth of astronomical literature poses significant challenges for researchers navigating and synthesizing general insights or even domain-specific knowledge. We present Pathfinder, a machine learning framework designed to enable literature review and knowledge discovery in astronomy, focusing on semantic searching with natural language instead of syntactic searches with keywords. Utilizing state-of-the-art large language models (LLMs) and a corpus of 350,000 peer-reviewed papers from the Astrophysics Data System (ADS), Pathfinder offers an innovative approach to scientific inquiry and literature exploration. Our framework couples advanced retrieval techniques with LLM-based synthesis to search astronomical literature by semantic context as a complement to currently existing methods that use keywords or citation graphs. It addresses complexities of jargon, named entities, and temporal aspects through time-based and citation-based weighting schemes. We demonstrate the tool's versatility through case studies, showcasing its application in various research scenarios. The system's performance is evaluated using custom benchmarks, including single-paper and multi-paper tasks. Beyond literature review, Pathfinder offers unique capabilities for reformatting answers in ways that are accessible to various audiences (e.g. in a different language or as simplified text), visualizing research landscapes, and tracking the impact of observatories and methodologies. This tool represents a significant advancement in applying AI to astronomical research, aiding researchers at all career stages in navigating modern astronomy literature.

astro-ph.IM

AstroLLaMA-Chat: Scaling AstroLLaMA with Conversational and Diverse Datasets

We explore the potential of enhancing LLM performance in astronomy-focused question-answering through targeted, continual pre-training. By employing a compact 7B-parameter LLaMA-2 model and focusing exclusively on a curated set of astronomy corpora -- comprising abstracts, introductions, and conclusions -- we achieve notable improvements in specialized topic comprehension. While general LLMs like GPT-4 excel in broader question-answering scenarios due to superior reasoning capabilities, our findings suggest that continual pre-training with limited resources can still enhance model performance on specialized topics. Additionally, we present an extension of AstroLLaMA: the fine-tuning of the 7B LLaMA model on a domain-specific conversational dataset, culminating in the release of the chat-enabled AstroLLaMA for community use. Comprehensive quantitative benchmarking is currently in progress and will be detailed in an upcoming full paper. The model, AstroLLaMA-Chat, is now available at https://huggingface.co/universeTBD, providing the first open-source conversational AI tool tailored for the astronomy community.

astro-ph.IM