SearcharxivSearch

arXiv subjects

Jeremy Kubica

Publications and source records attributed to Jeremy Kubica.

15 recordsLinked to original sources

Hyrax: An Extensible Framework for Rapid ML Experimentation and Unsupervised Discovery in the Era of Rubin, Roman, and Euclid

The NSF-DOE Vera C. Rubin Observatory, Roman Space Telescope, Euclid, and other next-generation surveys will deliver imaging, spectroscopic, and time-domain data at scales that increasingly shift the bottleneck in astronomical machine learning (ML) projects from model design to infrastructure. We present Hyrax, an open-source, modular, GPU-enabled Python framework that supports the full ML lifecycle in astronomy: from data acquisition and training to inference and experiment comparison, with capabilities including multimodal dataset support, integrated vector databases for similarity search, and interactive two- and three-dimensional latent-space exploration for unsupervised discovery. We demonstrate Hyrax's versatility through five representative applications on real survey data: (i) unsupervised representation learning on $\sim 4\times10^5$ Rubin Legacy Survey of Space and Time (LSST) Data Preview 1 (DP1) galaxies, surfacing new merger and low-surface-brightness candidates missing from reference Euclid and Dark Energy Survey catalogs, while also isolating imaging artifacts -- all without labeled training data; (ii) hybrid density-based clustering for identifying cluster-scale gravitational lens candidates in DP1 data; (iii) multimodal early-time transient classification in the Zwicky Transient Facility leveraging light curves, spectra, images, and metadata; (iv) supervised false-positive filtering in shift-and-stack searches for distant solar system objects in the Dark Energy Camera Ecliptic Exploration Project survey; and (v) supervised detection of semi-resolved dwarf galaxies in Hyper Suprime-Cam and LSST-like imaging using synthetic source injection. Together, these results demonstrate that Hyrax provides astronomy-specific ML infrastructure that enables systematic discovery and rapid methodological iteration across next-generation astronomical surveys.

astro-ph.IM

LightCurveLynx: Forward Modeling of Time-Domain Surveys with Application to ZTF SN Ia DR2

We present LightCurveLynx, a flexible and extensible software framework for end-to-end forward modeling time-domain light curves. Given the growing need for realistic simulations in the time-domain astronomy community, LightCurveLynx is designed to support a wide range of applications, including the development and validation of analysis pipelines, the optimization of survey strategies, and simulation-based inference studies. Realistic simulations can be generated from real survey metadata, forecasted survey plans, or user-defined mock survey strategies. We demonstrate the functionality of LightCurveLynx by generating a realistic simulation of Type Ia supernovae that is representative of the ZTF SN Ia Data Release 2 dataset and perform extensive comparisons between the simulated and observed samples to validate the software. The simulation shows excellent agreement with the data in parameter distributions (with the Kullback-Leibler divergence values around 0.01-0.02) and in noise properties. The Hubble diagram generated from the simulation also indicates that the sample is complete up to redshift 0.06, which is consistent with previous studies. Our results confirm that LightCurveLynx is robust, accurate, and ready for community use and contribution.

astro-ph.IM

Predictions of the LSST Solar System Yield: Neptune Trojans

The NSF-DOE Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST), beginning full operations in late 2025, will dramatically transform solar system science by vastly expanding discoveries and providing detailed characterization opportunities across all small body populations. This includes the co-orbiting 1:1 resonant Neptune Trojans, which are thought to be dynamically hot captures from the protoplanetary disk. Using the survey simulator $\texttt{Sorcha}$, combined with the latest LSST cadence simulations, we present the very first predictions for the Neptune Trojan yield within the LSST. We forecast a model-dependent median number of $\sim130-300$ discovered Neptune Trojans, and infer a notable 2:1 detection bias toward the recently emerged L5 cloud near the galactic plane versus the L4 cloud, reflecting the lower-cadence coverage in the Northern Ecliptic Spur region that suppresses L4 detections. The additionally simulated Science Validation survey will offer the very first early insights into this understudied cloud. Around 60\% of detected main survey Neptune Trojans will meet stringent color light curve quality criteria, increasing the sample size more than fourfold compared to existing datasets. This enhanced sample will enable robust statistical analyses of Neptune Trojan color and size distributions, crucial for understanding their origins and relationship to the broader trans-Neptunian population. These comprehensive color measurements represent a major step forward in characterizing the Neptune Trojan population and will facilitate future targeted spectroscopic observations.

astro-ph.EP

Using LSDB to enable large-scale catalog distribution, cross-matching, and analytics

The Vera C. Rubin Observatory will generate an unprecedented volume of data, including approximately 60 petabytes of raw data and around 30 trillion observed sources, posing a significant challenge for large-scale and end-user scientific analysis. As part of the LINCC Frameworks Project we are addressing these challenges with the development of the HATS (Hierarchical Adaptive Tiling Scheme) format and analysis package LSDB. HATS partitions data adaptively using a hierarchical tiling system to balance the file sizes, enabling efficient parallel analysis. Recent updates include improved metadata consistency, support for incremental updates, and enhanced compatibility with evolving datasets. LSDB complements HATS by providing a scalable, user-friendly interface for large catalog analysis, integrating spatial queries, crossmatching, and time-series tools while utilizing Dask for parallelization. We have successfully demonstrated the use of these tools with datasets such as ZTF and Pan-STARRS data releases on both cluster and cloud environments. We are deeply involved in several ongoing collaborations to ensure alignment with community needs, with future plans for IVOA standardization and support for upcoming Rubin, Euclid and Roman data. We provide our code and materials at lsdb.io.

astro-ph.IM

Variability-finding in Rubin Data Preview 1 with LSDB

The Vera C. Rubin Observatory recently released Data Preview 1 (DP1) in advance of the upcoming Legacy Survey of Space and Time (LSST), which will enable boundless discoveries in time-domain astronomy over the next ten years. DP1 provides an ideal sandbox for validating innovative data analysis approaches for the LSST mission, whose scale challenges established software infrastructure paradigms. This note presents a pair of such pipelines for variability-finding using powerful software infrastructure suited to LSST data, namely the HATS (Hierarchical Adaptive Tiling Scheme) format and the LSDB framework, developed by the LSST Interdisciplinary Network for Collaboration and Computing (LINCC) Frameworks team. This article presents a pair of variability-finding pipelines built on LSDB, the HATS catalog of DP1 data, and preliminary results of detected variable objects, two of which are novel discoveries.

astro-ph.IM

Predictions of the LSST Solar System Yield: Near-Earth Objects, Main Belt Asteroids, Jupiter Trojans, and Trans-Neptunian Objects

The NSF-DOE Vera C. Rubin Observatory is a new 8m-class survey facility presently being commissioned in Chile, expected to begin the 10yr-long Legacy Survey of Space and Time (LSST) by the end of 2025. Using the purpose-built Sorcha survey simulator (Merritt et al. In Press), and near-final observing cadence, we perform the first high-fidelity simulation of LSST's solar system catalog for key small body populations. We show that the final LSST catalog will deliver over 1.1 billion observations of small bodies and raise the number of known objects to 1.27E5 near-Earth objects, 5.09E6 main belt asteroids, 1.09E5 Jupiter Trojans, and 3.70E4 trans-Neptunian objects. These represent 4-9x more objects than are presently known in each class, making LSST the largest source of data for small body science in this and the following decade. We characterize the measurements available for these populations, including orbits, griz colors, and lightcurves, and point out science opportunities they open. Importantly, we show that ~70% of the main asteroid belt and more distant populations will be discovered in the first two years of the survey, making high-impact solar system science possible from very early on. We make our simulated LSST catalog publicly available, allowing researchers to test their methods on an up-to-date, representative, full-scale simulation of LSST data.

astro-ph.EP

Predictions of the LSST Solar System Yield: Discovery Rates and Characterizations of Centaurs

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will start by the end of 2025 and operate for ten years, offering billions of observations of the southern night sky. One of its main science goals is to create an inventory of the Solar System, allowing for a more detailed understanding of small body populations including the Centaurs, which will benefit from the survey's high cadence and depth. In this paper, we establish the first discovery limits for Centaurs throughout the LSST's decade-long operation using the best available dynamical models. Using the survey simulator $\texttt{Sorcha}$, we predict a $\sim$7-12 fold increase in Centaurs in the Minor Planet Center (MPC) database, reaching $\sim$1200-2000 (dependent on definition) by the end of the survey - about 50$\%$ of which are expected within the first 2 years. Approximately 30-50 Centaurs will be observed twice as frequently as they fall within one of the LSST's Deep Drilling Fields (DDF) for on average only up to two months. Outside of the DDFs, Centaurs will receive $\sim$200 observations across the $\textit{ugrizy}$ filter range, facilitating searches for cometary-like activity through PSF extension analysis, as well as fitting light-curves and phase curves for color determination. Regardless of definition, over 200 Centaurs will achieve high-quality color measurements across at least three filters in the LSST's six filters. These observations will also provide over 300 well-defined phase curves in the $\textit{griz}$ bands, improving absolute magnitude measurements to a precision of 0.2 mags.

astro-ph.EP

Sorcha: A Solar System Survey Simulator for the Legacy Survey of Space and Time

The upcoming Legacy Survey of Space and Time (LSST) at the Vera C. Rubin Observatory is expected to revolutionize solar system astronomy. Unprecedented in scale, this ten-year wide-field survey will collect billions of observations and discover a predicted $\sim$5 million new solar system objects. Like all astronomical surveys, its results will be affected by a complex system of intertwined detection biases. Survey simulators have long been used to forward-model the effects of these biases on a given population, allowing for a direct comparison to real discoveries. However, the scale and tremendous scope of the LSST requires the development of new tools. In this paper we present Sorcha, an open-source survey simulator written in Python. Designed with the scale of LSST in mind, Sorcha is a comprehensive survey simulator to cover all solar system small-body populations. Its flexible, modular design allows Sorcha to be easily adapted to other surveys by the user. The simulator is built to run both locally and on high-performance computing (HPC) clusters, allowing for repeated simulation of millions to billions of objects (both real and synthetic).

astro-ph.EP

Sorcha: Optimized Solar System Ephemeris Generation

Sorcha is a solar system survey simulator built for the Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) and future large-scale wide-field surveys. Over the ten-year survey, the LSST is expected to collect roughly a billion observations of minor planets. The task of a solar system survey simulator is to take a set of input objects (described by orbits and physical properties) and determine what a real or hypothetical survey would have discovered. Existing survey simulators have a computational bottleneck in determining which input objects lie in each survey field, making them infeasible for LSST data scales. Sorcha can swiftly, efficiently, and accurately calculate the on-sky positions for sets of millions of input orbits and surveys with millions of visits, identifying which exposures these objects cross, in order for later stages of the software to make detailed estimates of the apparent magnitude and detectability of those input small bodies. In this paper, we provide the full details of the algorithm and software behind Sorcha's ephemeris generator. Like many of Sorcha's components, its ephemeris generator can be easily used for other surveys.

astro-ph.EP

DeepDISC-photoz: Deep Learning-Based Photometric Redshift Estimation for Rubin LSST

Photometric redshifts will be a key data product for the Rubin Observatory Legacy Survey of Space and Time (LSST) as well as for future ground and space-based surveys. The need for photometric redshifts, or photo-zs, arises from sparse spectroscopic coverage of observed galaxies. LSST is expected to observe billions of objects, making it crucial to have a photo-z estimator that is accurate and efficient. To that end, we present DeepDISC photo-z, a photo-z estimator that is an extension of the DeepDISC framework. The base DeepDISC network simultaneously detects, segments, and classifies objects in multi-band coadded images. We introduce photo-z capabilities to DeepDISC by adding a redshift estimation Region of Interest head, which produces a photo-z probability distribution function for each detected object. On simulated LSST images, DeepDISC photo-z outperforms traditional catalog-based estimators, in both point estimate and probabilistic metrics. We validate DeepDISC by examining dependencies on systematics including galactic extinction, blending and PSF effects. We also examine the impact of the data quality and the size of the training set and model. We find that the biggest factor in DeepDISC photo-z quality is the signal-to-noise of the imaging data, and see a reduction in photo-z scatter approximately proportional to the image data signal-to-noise. Our code is fully public and integrated in the RAIL photo-z package for ease of use and comparison to other codes at https://github.com/LSSTDESC/rail_deepdisc

astro-ph.IM

Superphot+: Realtime Fitting and Classification of Supernova Light Curves

Photometric classifications of supernova (SN) light curves have become necessary to utilize the full potential of large samples of observations obtained from wide-field photometric surveys, such as the Zwicky Transient Facility (ZTF) and the Vera C. Rubin Observatory. Here, we present a photometric classifier for SN light curves that does not rely on redshift information and still maintains comparable accuracy to redshift-dependent classifiers. Our new package, Superphot+, uses a parametric model to extract meaningful features from multiband SN light curves. We train a gradient-boosted machine with fit parameters from 6,061 ZTF SNe that pass data quality cuts and are spectroscopically classified as one of five classes: SN Ia, SN II, SN Ib/c, SN IIn, and SLSN-I. Without redshift information, our classifier yields a class-averaged F1-score of 0.61 +/- 0.02 and a total accuracy of 0.83 +/- 0.01. Including redshift information improves these metrics to 0.71 +/- 0.02 and 0.88 +/- 0.01, respectively. We assign new class probabilities to 3,558 ZTF transients that show SN-like characteristics (based on the ALeRCE Broker light curve and stamp classifiers), but lack spectroscopic classifications. Finally, we compare our predicted SN labels with those generated by the ALeRCE light curve classifier, finding that the two classifiers agree on photometric labels for 82 +/- 2% of light curves with spectroscopic labels and 72% of light curves without spectroscopic labels. Superphot+ is currently classifying ZTF SNe in real time via the ANTARES Broker, and is designed for simple adaptation to six-band Rubin light curves in the future.

astro-ph.HE

From Data to Software to Science with the Rubin Observatory LSST

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) dataset will dramatically alter our understanding of the Universe, from the origins of the Solar System to the nature of dark matter and dark energy. Much of this research will depend on the existence of robust, tested, and scalable algorithms, software, and services. Identifying and developing such tools ahead of time has the potential to significantly accelerate the delivery of early science from LSST. Developing these collaboratively, and making them broadly available, can enable more inclusive and equitable collaboration on LSST science. To facilitate such opportunities, a community workshop entitled "From Data to Software to Science with the Rubin Observatory LSST" was organized by the LSST Interdisciplinary Network for Collaboration and Computing (LINCC) and partners, and held at the Flatiron Institute in New York, March 28-30th 2022. The workshop included over 50 in-person attendees invited from over 300 applications. It identified seven key software areas of need: (i) scalable cross-matching and distributed joining of catalogs, (ii) robust photometric redshift determination, (iii) software for determination of selection functions, (iv) frameworks for scalable time-series analyses, (v) services for image access and reprocessing at scale, (vi) object image access (cutouts) and analysis at scale, and (vii) scalable job execution systems. This white paper summarizes the discussions of this workshop. It considers the motivating science use cases, identified cross-cutting algorithms, software, and services, their high-level technical specifications, and the principles of inclusive collaborations needed to develop them. We provide it as a useful roadmap of needs, as well as to spur action and collaboration between groups and individuals looking to develop reusable software for early LSST science.

astro-ph.IM

Algorithms and Statistical Models for Scientific Discovery in the Petabyte Era

The field of astronomy has arrived at a turning point in terms of size and complexity of both datasets and scientific collaboration. Commensurately, algorithms and statistical models have begun to adapt --- e.g., via the onset of artificial intelligence --- which itself presents new challenges and opportunities for growth. This white paper aims to offer guidance and ideas for how we can evolve our technical and collaborative frameworks to promote efficient algorithmic development and take advantage of opportunities for scientific discovery in the petabyte era. We discuss challenges for discovery in large and complex data sets; challenges and requirements for the next stage of development of statistical methodologies and algorithmic tool sets; how we might change our paradigms of collaboration and education; and the ethical implications of scientists' contributions to widely applicable algorithms and computational modeling. We start with six distinct recommendations that are supported by the commentary following them. This white paper is related to a larger corpus of effort that has taken place within and around the Petabytes to Science Workshops (https://petabytestoscience.github.io/).

astro-ph.IM

The Pan-STARRS Moving Object Processing System

We describe the Pan-STARRS Moving Object Processing System (MOPS), a modern software package that produces automatic asteroid discoveries and identifications from catalogs of transient detections from next-generation astronomical survey telescopes. MOPS achieves > 99.5% efficiency in producing orbits from a synthetic but realistic population of asteroids whose measurements were simulated for a Pan-STARRS4-class telescope. Additionally, using a non-physical grid population, we demonstrate that MOPS can detect populations of currently unknown objects such as interstellar asteroids. MOPS has been adapted successfully to the prototype Pan-STARRS1 telescope despite differences in expected false detection rates, fill-factor loss and relatively sparse observing cadence compared to a hypothetical Pan-STARRS4 telescope and survey. MOPS remains >99.5% efficient at detecting objects on a single night but drops to 80% efficiency at producing orbits for objects detected on multiple nights. This loss is primarily due to configurable MOPS processing limits that are not yet tuned for the Pan-STARRS1 mission. The core MOPS software package is the product of more than 15 person-years of software development and incorporates countless additional years of effort in third-party software to perform lower-level functions such as spatial searching or orbit determination. We describe the high-level design of MOPS and essential subcomponents, the suitability of MOPS for other survey programs, and suggest a road map for future MOPS development.

astro-ph.IM

Efficient intra- and inter-night linking of asteroid detections using kd-trees

The Panoramic Survey Telescope And Rapid Response System (Pan-STARRS) under development at the University of Hawaii's Institute for Astronomy is creating the first fully automated end-to-end Moving Object Processing System (MOPS) in the world. It will be capable of identifying detections of moving objects in our solar system and linking those detections within and between nights, attributing those detections to known objects, calculating initial and differentially-corrected orbits for linked detections, precovering detections when they exist, and orbit identification. Here we describe new kd-tree and variable-tree algorithms that allow fast, efficient, scalable linking of intra and inter-night detections. Using a pseudo-realistic simulation of the Pan-STARRS survey strategy incorporating weather, astrometric accuracy and false detections we have achieved nearly 100% efficiency and accuracy for intra-night linking and nearly 100% efficiency for inter-night linking within a lunation. At realistic sky-plane densities for both real and false detections the intra-night linking of detections into `tracks' currently has an accuracy of 0.3%. Successful tests of the MOPS on real source detections from the Spacewatch asteroid survey indicate that the MOPS is capable of identifying asteroids in real data.

astro-ph