SearcharxivSearch

arXiv subjects

Drew Oldag

Publications and source records attributed to Drew Oldag.

10 recordsLinked to original sources

Catching Disguised Transients with ASTRANet: Anomaly-Aware Spectroscopic Classification and Conformal Calibration

Time-domain surveys discover thousands of transients per year, but the spectroscopic identification of rare and physically peculiar objects remains rate-limited by closed-set classifiers that confidently assign every input to a known class -- including spectra that genuinely belong to no known class. We present the \texttt{ASTRANet} framework, a confidence-aware infrastructure for spectroscopic transient classification built around three coupled modules: a hierarchical spectral classifier that operates directly on observer-frame spectra without requiring host-galaxy redshift or spectral phase as inputs; an anomaly detection layer (\texttt{ASTRANet-Sentinel}) that non-linearly combines $16$ embedding-space anomaly scores spanning four physically motivated families; and a conformal uncertainty quantification layer (\texttt{ASTRANet-CP}). We validate the framework on a held-out evaluation set of $289$ rare and out-of-taxonomy transients spanning $11$ classes deliberately excluded from training, chosen to span the full physical diversity of the rare-anomaly population: AGN-related outliers, GRB-related events, gap transients, novae, and peculiar supernovae. Through five astrophysically distinct failure modes of closed-set classifiers, we show that classifier-internal uncertainty and embedding-based anomaly detection are structurally complementary axes of confidence rather than alternative implementations of the same estimator. We further introduce AD-stratified Mondrian conformal prediction (AD-MCP) within \texttt{ASTRANet-CP}, achieving uniform conditional coverage across anomaly-score strata where vanilla Mondrian under-covers in the operational regime. This establishes the methodological infrastructure for confidence-aware spectroscopic discovery in the Vera C.\ Rubin Observatory era.

astro-ph.IM

Hyrax: An Extensible Framework for Rapid ML Experimentation and Unsupervised Discovery in the Era of Rubin, Roman, and Euclid

The NSF-DOE Vera C. Rubin Observatory, Roman Space Telescope, Euclid, and other next-generation surveys will deliver imaging, spectroscopic, and time-domain data at scales that increasingly shift the bottleneck in astronomical machine learning (ML) projects from model design to infrastructure. We present Hyrax, an open-source, modular, GPU-enabled Python framework that supports the full ML lifecycle in astronomy: from data acquisition and training to inference and experiment comparison, with capabilities including multimodal dataset support, integrated vector databases for similarity search, and interactive two- and three-dimensional latent-space exploration for unsupervised discovery. We demonstrate Hyrax's versatility through five representative applications on real survey data: (i) unsupervised representation learning on $\sim 4\times10^5$ Rubin Legacy Survey of Space and Time (LSST) Data Preview 1 (DP1) galaxies, surfacing new merger and low-surface-brightness candidates missing from reference Euclid and Dark Energy Survey catalogs, while also isolating imaging artifacts -- all without labeled training data; (ii) hybrid density-based clustering for identifying cluster-scale gravitational lens candidates in DP1 data; (iii) multimodal early-time transient classification in the Zwicky Transient Facility leveraging light curves, spectra, images, and metadata; (iv) supervised false-positive filtering in shift-and-stack searches for distant solar system objects in the Dark Energy Camera Ecliptic Exploration Project survey; and (v) supervised detection of semi-resolved dwarf galaxies in Hyper Suprime-Cam and LSST-like imaging using synthetic source injection. Together, these results demonstrate that Hyrax provides astronomy-specific ML infrastructure that enables systematic discovery and rapid methodological iteration across next-generation astronomical surveys.

astro-ph.IM

Diagnosing the Effects of Spectroscopic Training Set Imperfection on Photometric Redshift Performance

Most LSST extragalactic science will rely on photometric redshifts (photo-$z$) to extract distance information for the galaxies. However, an incomplete or non-representative training set can introduce bias into photo-$z$ estimation. It is necessary to understand how various forms of training set imperfection, such as incompleteness and non-trivial spectroscopic target selection, affect photo-$z$ estimation algorithms, and to identify metrics best-suited to quantify the impact. This work aims to systematically study metrics for diagnosing how various photo-$z$ methods react to certain types of training set incompleteness and non-representativeness. We use methods available through the open-source Python library Redshift Assessment Infrastructure Layers (RAIL) to systematically test the algorithms CMNN, GPz, FlexZBoost, and PZFlow on mock training data degraded in accordance with several existing spectroscopic sky surveys, as well as under conditions of inverse redshift incompleteness, which approximately mimics observed patterns of incompleteness at high redshift. We employ the algorithm TrainZ as a control. Finally, we quantify photo-$z$ algorithm performance using a variety of statistical metrics implemented externally to RAIL. We determine that the Kullback-Liebler Divergence, Wasserstein Distance, and Probability Integral Transform are particularly informative metrics with which to assess the impact of training set imperfection on algorithmic performance. We also find that inverse redshift incompleteness effects alone lack the complexity to realistically represent anticipated training data.

astro-ph.IM

Predictions of the LSST Solar System Yield: Neptune Trojans

The NSF-DOE Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST), beginning full operations in late 2025, will dramatically transform solar system science by vastly expanding discoveries and providing detailed characterization opportunities across all small body populations. This includes the co-orbiting 1:1 resonant Neptune Trojans, which are thought to be dynamically hot captures from the protoplanetary disk. Using the survey simulator $\texttt{Sorcha}$, combined with the latest LSST cadence simulations, we present the very first predictions for the Neptune Trojan yield within the LSST. We forecast a model-dependent median number of $\sim130-300$ discovered Neptune Trojans, and infer a notable 2:1 detection bias toward the recently emerged L5 cloud near the galactic plane versus the L4 cloud, reflecting the lower-cadence coverage in the Northern Ecliptic Spur region that suppresses L4 detections. The additionally simulated Science Validation survey will offer the very first early insights into this understudied cloud. Around 60\% of detected main survey Neptune Trojans will meet stringent color light curve quality criteria, increasing the sample size more than fourfold compared to existing datasets. This enhanced sample will enable robust statistical analyses of Neptune Trojan color and size distributions, crucial for understanding their origins and relationship to the broader trans-Neptunian population. These comprehensive color measurements represent a major step forward in characterizing the Neptune Trojan population and will facilitate future targeted spectroscopic observations.

astro-ph.EP

Predictions of the LSST Solar System Yield: Discovery Rates and Characterizations of Centaurs

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will start by the end of 2025 and operate for ten years, offering billions of observations of the southern night sky. One of its main science goals is to create an inventory of the Solar System, allowing for a more detailed understanding of small body populations including the Centaurs, which will benefit from the survey's high cadence and depth. In this paper, we establish the first discovery limits for Centaurs throughout the LSST's decade-long operation using the best available dynamical models. Using the survey simulator $\texttt{Sorcha}$, we predict a $\sim$7-12 fold increase in Centaurs in the Minor Planet Center (MPC) database, reaching $\sim$1200-2000 (dependent on definition) by the end of the survey - about 50$\%$ of which are expected within the first 2 years. Approximately 30-50 Centaurs will be observed twice as frequently as they fall within one of the LSST's Deep Drilling Fields (DDF) for on average only up to two months. Outside of the DDFs, Centaurs will receive $\sim$200 observations across the $\textit{ugrizy}$ filter range, facilitating searches for cometary-like activity through PSF extension analysis, as well as fitting light-curves and phase curves for color determination. Regardless of definition, over 200 Centaurs will achieve high-quality color measurements across at least three filters in the LSST's six filters. These observations will also provide over 300 well-defined phase curves in the $\textit{griz}$ bands, improving absolute magnitude measurements to a precision of 0.2 mags.

astro-ph.EP

Sorcha: A Solar System Survey Simulator for the Legacy Survey of Space and Time

The upcoming Legacy Survey of Space and Time (LSST) at the Vera C. Rubin Observatory is expected to revolutionize solar system astronomy. Unprecedented in scale, this ten-year wide-field survey will collect billions of observations and discover a predicted $\sim$5 million new solar system objects. Like all astronomical surveys, its results will be affected by a complex system of intertwined detection biases. Survey simulators have long been used to forward-model the effects of these biases on a given population, allowing for a direct comparison to real discoveries. However, the scale and tremendous scope of the LSST requires the development of new tools. In this paper we present Sorcha, an open-source survey simulator written in Python. Designed with the scale of LSST in mind, Sorcha is a comprehensive survey simulator to cover all solar system small-body populations. Its flexible, modular design allows Sorcha to be easily adapted to other surveys by the user. The simulator is built to run both locally and on high-performance computing (HPC) clusters, allowing for repeated simulation of millions to billions of objects (both real and synthetic).

astro-ph.EP

Predictions of the LSST Solar System Yield: Near-Earth Objects, Main Belt Asteroids, Jupiter Trojans, and Trans-Neptunian Objects

The NSF-DOE Vera C. Rubin Observatory is a new 8m-class survey facility presently being commissioned in Chile, expected to begin the 10yr-long Legacy Survey of Space and Time (LSST) by the end of 2025. Using the purpose-built Sorcha survey simulator (Merritt et al. In Press), and near-final observing cadence, we perform the first high-fidelity simulation of LSST's solar system catalog for key small body populations. We show that the final LSST catalog will deliver over 1.1 billion observations of small bodies and raise the number of known objects to 1.27E5 near-Earth objects, 5.09E6 main belt asteroids, 1.09E5 Jupiter Trojans, and 3.70E4 trans-Neptunian objects. These represent 4-9x more objects than are presently known in each class, making LSST the largest source of data for small body science in this and the following decade. We characterize the measurements available for these populations, including orbits, griz colors, and lightcurves, and point out science opportunities they open. Importantly, we show that ~70% of the main asteroid belt and more distant populations will be discovered in the first two years of the survey, making high-impact solar system science possible from very early on. We make our simulated LSST catalog publicly available, allowing researchers to test their methods on an up-to-date, representative, full-scale simulation of LSST data.

astro-ph.EP

Sorcha: Optimized Solar System Ephemeris Generation

Sorcha is a solar system survey simulator built for the Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) and future large-scale wide-field surveys. Over the ten-year survey, the LSST is expected to collect roughly a billion observations of minor planets. The task of a solar system survey simulator is to take a set of input objects (described by orbits and physical properties) and determine what a real or hypothetical survey would have discovered. Existing survey simulators have a computational bottleneck in determining which input objects lie in each survey field, making them infeasible for LSST data scales. Sorcha can swiftly, efficiently, and accurately calculate the on-sky positions for sets of millions of input orbits and surveys with millions of visits, identifying which exposures these objects cross, in order for later stages of the software to make detailed estimates of the apparent magnitude and detectability of those input small bodies. In this paper, we provide the full details of the algorithm and software behind Sorcha's ephemeris generator. Like many of Sorcha's components, its ephemeris generator can be easily used for other surveys.

astro-ph.EP

Redshift Assessment Infrastructure Layers (RAIL): Rubin-era photometric redshift stress-testing and at-scale production

Virtually all extragalactic use cases of the Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) require the use of galaxy redshift information, yet the vast majority of its sample of tens of billions of galaxies will lack high-fidelity spectroscopic measurements thereof, instead relying on photometric redshifts (photo-$z$) subject to systematic imprecision and inaccuracy best encapsulated by photo-$z$ probability density functions (PDFs). We present the version 1 release of Redshift Assessment Infrastructure Layers (RAIL), an open source Python library for at-scale probabilistic photo-$z$ estimation, initiated by the LSST Dark Energy Science Collaboration (DESC) with contributions from the LSST Interdisciplinary Network for Collaboration and Computing (LINCC) Frameworks team. RAIL's three subpackages provide modular tools for end-to-end stress-testing, including a forward modeling suite to generate realistically complex photometry, a unified API for estimating per-galaxy and ensemble redshift PDFs by an extensible set of algorithms, and built-in metrics of both photo-$z$ PDFs and point estimates. RAIL serves as a flexible toolkit enabling the derivation and optimization of photo-$z$ data products at scale for a variety of science goals and is not specific to LSST data. We thus describe to the extragalactic science community, including and beyond Rubin the design and functionality of the RAIL software library so that any researcher may have access to its wide array of photo-$z$ characterization and assessment tools.

astro-ph.IM

DeepDISC-photoz: Deep Learning-Based Photometric Redshift Estimation for Rubin LSST

Photometric redshifts will be a key data product for the Rubin Observatory Legacy Survey of Space and Time (LSST) as well as for future ground and space-based surveys. The need for photometric redshifts, or photo-zs, arises from sparse spectroscopic coverage of observed galaxies. LSST is expected to observe billions of objects, making it crucial to have a photo-z estimator that is accurate and efficient. To that end, we present DeepDISC photo-z, a photo-z estimator that is an extension of the DeepDISC framework. The base DeepDISC network simultaneously detects, segments, and classifies objects in multi-band coadded images. We introduce photo-z capabilities to DeepDISC by adding a redshift estimation Region of Interest head, which produces a photo-z probability distribution function for each detected object. On simulated LSST images, DeepDISC photo-z outperforms traditional catalog-based estimators, in both point estimate and probabilistic metrics. We validate DeepDISC by examining dependencies on systematics including galactic extinction, blending and PSF effects. We also examine the impact of the data quality and the size of the training set and model. We find that the biggest factor in DeepDISC photo-z quality is the signal-to-noise of the imaging data, and see a reduction in photo-z scatter approximately proportional to the image data signal-to-noise. Our code is fully public and integrated in the RAIL photo-z package for ease of use and comparison to other codes at https://github.com/LSSTDESC/rail_deepdisc

astro-ph.IM