SearcharxivSearch

arXiv subjects

Emmanuel Gangler

Publications and source records attributed to Emmanuel Gangler.

15 recordsLinked to original sources

SNAD: enabling discovery in the era of big data

In the era of wide-field surveys and big data in astronomy, the SNAD team is exploiting the potential of modern datasets for discovering new, unforeseen, or rare astrophysical objects and phenomena with machine learning (ML). The SNAD pipeline was built under the hypothesis that, although automatic ML algorithms have a crucial role to play in this task, the scientific discovery is only completely realized when such systems are designed to boost the impact of domain knowledge experts. Our key contributions include the development of the Coniferest Python library, which offers implementations of two active learning algorithms with an ``expert in loop'', and the creation of the SNAD Transient Miner, facilitating the search for specific types of transients. We have also developed the SNAD Viewer, a web portal that provides a centralized view of individual objects from the Zwicky Transient Facility's (ZTF) data releases, making the analysis of potential anomalies more efficient. Finally, when applied to ZTF data, our approach has resulted in more than a hundred new supernova (SN) candidates, along with a few other non-catalogued objects, such as red dwarf flares, superluminous SNe, RS CVn type variables, and young stellar objects.

astro-ph.HE

The Vera C. Rubin Observatory Data Preview 1

We present Rubin Data Preview 1 DP1, the first data from the NSF DOE Vera C Rubin Observatory, comprising raw and calibrated single epoch images, coadds, difference images, detection catalogs, and ancillary data products. DP1 is based on 1792 optical near infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera LSSTComCam on the Simonyi Survey Telescope at the Summit Facility on Cerro Pach\'on Chile in late 2024. DP1 covers $\sim$15 deg$^2$ distributed across seven roughly equal-sized non-contiguous fields, each independently observed in six broad photometric bands $ugrizy$. The median FWHM of the point spread function across all bands is approximately 1.14 arcseconds, with the sharpest images reaching about 0.58 arcseconds. The 5$\sigma$ point source depths for coadded images in the deepest field the Extended Chandra Deep Field South are $u$ = 24.55, $g$ = 26.18, $r$ = 25.96, $i$ = 25.71, $z$ = 25.07, $y$ = 23.1. Other fields are no more than 2.2 magnitudes shallower in any band where they have nonzero coverage. DP1 contains approximately 2.3 million distinct astrophysical objects, of which 1.6 million are extended in at least one band in coadds and 431 solar system objects of which 93 are new discoveries. DP1 is approximately 3.5 TB in size and is available to Rubin data rights holders via the Rubin Science Platform a cloud based environment for the analysis of petascale astronomical data. While small compared to future LSST releases its high quality and diversity of data support a broad range of early science investigations ahead of full operations in 2026.

astro-ph.IM

What ZTF Saw Where Rubin Looked: Anomaly Hunting in DR23

We present results from the SNAD VIII Workshop, during which we conducted the first systematic anomaly search in the ZTF fields also observed by LSSTComCam during Rubin Scientific Pipeline commissioning. Using the PineForest active anomaly detection algorithm, we analysed four selected fields (two galactic and two extragalactic) and visually inspected 400 candidates. As a result, we discovered six previously uncatalogued variable stars, including RS~CVn, BY Draconis, ellipsoidal, and solar-type variables, and refined classifications and periods for six known objects. These results demonstrate the effectiveness of the SNAD anomaly detection pipeline and provide a preview of the discovery potential in the upcoming LSST data.

astro-ph.IM

Variability-finding in Rubin Data Preview 1 with LSDB

The Vera C. Rubin Observatory recently released Data Preview 1 (DP1) in advance of the upcoming Legacy Survey of Space and Time (LSST), which will enable boundless discoveries in time-domain astronomy over the next ten years. DP1 provides an ideal sandbox for validating innovative data analysis approaches for the LSST mission, whose scale challenges established software infrastructure paradigms. This note presents a pair of such pipelines for variability-finding using powerful software infrastructure suited to LSST data, namely the HATS (Hierarchical Adaptive Tiling Scheme) format and the LSDB framework, developed by the LSST Interdisciplinary Network for Collaboration and Computing (LINCC) Frameworks team. This article presents a pair of variability-finding pipelines built on LSDB, the HATS catalog of DP1 data, and preliminary results of detected variable objects, two of which are novel discoveries.

astro-ph.IM

Signatures to help interpretability of anomalies

Machine learning is often viewed as a black box when it comes to understanding its output, be it a decision or a score. Automatic anomaly detection is no exception to this rule, and quite often the astronomer is left to independently analyze the data in order to understand why a given event is tagged as an anomaly. We introduce here idea of anomaly signature, whose aim is to help the interpretability of anomalies by highlighting which features contributed to the decision.

cs.LG

Dataset of artefacts for machine learning applications in astronomy

Accurate photometry in astronomical surveys is challenged by image artefacts, which affect measurements and degrade data quality. Due to the large amount of available data, this task is increasingly handled using machine learning algorithms, which often require a labelled training set to learn data patterns. We present an expert-labelled dataset of 1127 artefacts with 1213 labels from 26 fields in ZTF DR3, along with a complementary set of nominal objects. The artefact dataset was compiled using the active anomaly detection algorithm PineForest, developed by the SNAD team. These datasets can serve as valuable resources for real-bogus classification, catalogue cleaning, anomaly detection, and educational purposes. Both artefacts and nominal images are provided in FITS format in two sizes (28 x 28 and 63 x 63 pixels). The datasets are publicly available for further scientific applications.

astro-ph.IM

Exploring the Universe with SNAD: Anomaly Detection in Astronomy

SNAD is an international project with a primary focus on detecting astronomical anomalies within large-scale surveys, using active learning and other machine learning algorithms. The work carried out by SNAD not only contributes to the discovery and classification of various astronomical phenomena but also enhances our understanding and implementation of machine learning techniques within the field of astrophysics. This paper provides a review of the SNAD project and summarizes the advancements and achievements made by the team over several years.

astro-ph.IM

Multi-View Symbolic Regression

Symbolic regression (SR) searches for analytical expressions representing the relationship between a set of explanatory and response variables. Current SR methods assume a single dataset extracted from a single experiment. Nevertheless, frequently, the researcher is confronted with multiple sets of results obtained from experiments conducted with different setups. Traditional SR methods may fail to find the underlying expression since the parameters of each experiment can be different. In this work we present Multi-View Symbolic Regression (MvSR), which takes into account multiple datasets simultaneously, mimicking experimental environments, and outputs a general parametric solution. This approach fits the evaluated expression to each independent dataset and returns a parametric family of functions f(x; theta) simultaneously capable of accurately fitting all datasets. We demonstrate the effectiveness of MvSR using data generated from known expressions, as well as real-world data from astronomy, chemistry and economy, for which an a priori analytical expression is not available. Results show that MvSR obtains the correct expression more frequently and is robust to hyperparameters change. In real-world data, it is able to grasp the group behavior, recovering known expressions from the literature as well as promising alternatives, thus enabling the use of SR to a large range of experimental scenarios.

cs.LG

Explainable classification of astronomical uncertain time series

Exploring the expansion history of the universe, understanding its evolutionary stages, and predicting its future evolution are important goals in astrophysics. Today, machine learning tools are used to help achieving these goals by analyzing transient sources, which are modeled as uncertain time series. Although black-box methods achieve appreciable performance, existing interpretable time series methods failed to obtain acceptable performance for this type of data. Furthermore, data uncertainty is rarely taken into account in these methods. In this work, we propose an uncertaintyaware subsequence based model which achieves a classification comparable to that of state-of-the-art methods. Unlike conformal learning which estimates model uncertainty on predictions, our method takes data uncertainty as additional input. Moreover, our approach is explainable-by-design, giving domain experts the ability to inspect the model and explain its predictions. The explainability of the proposed method has also the potential to inspire new developments in theoretical astrophysics modeling by suggesting important subsequences which depict details of light curve shapes. The dataset, the source code of our experiment, and the results are made available on a public repository.

cs.LG

Supernova search with active learning in ZTF DR3

We provide the first results from the complete SNAD adaptive learning pipeline in the context of a broad scope of data from large-scale astronomical surveys. The main goal of this work is to explore the potential of adaptive learning techniques in application to big data sets. Our SNAD team used Active Anomaly Discovery (AAD) as a tool to search for new supernova (SN) candidates in the photometric data from the first 9.4 months of the Zwicky Transient Facility (ZTF) survey, namely, between March 17 and December 31 2018 (58194 < MJD < 58483). We analysed 70 ZTF fields at a high galactic latitude and visually inspected 2100 outliers. This resulted in 104 SN-like objects being found, 57 of which were reported to the Transient Name Server for the first time and with 47 having previously been mentioned in other catalogues, either as SNe with known types or as SN candidates. We visually inspected the multi-colour light curves of the non-catalogued transients and performed fittings with different supernova models to assign it to a probable photometric class: Ia, Ib/c, IIP, IIL, or IIn. Moreover, we also identified unreported slow-evolving transients that are good superluminous SN candidates, along with a few other non-catalogued objects, such as red dwarf flares and active galactic nuclei. Beyond confirming the effectiveness of human-machine integration underlying the AAD strategy, our results shed light on potential leaks in currently available pipelines. These findings can help avoid similar losses in future large-scale astronomical surveys. Furthermore, the algorithm enables direct searches of any type of data and based on any definition of an anomaly set by the expert.

astro-ph.HE

Gravitation And the Universe from large Scale-Structures: The GAUSS mission concept

Today, thanks in particular to the results of the ESA Planck mission, the concordance cosmological model appears to be the most robust to describe the evolution and content of the Universe from its early to late times. It summarizes the evolution of matter, made mainly of dark matter, from the primordial fluctuations generated by inflation around $10^{-30}$ second after the Big-Bang to galaxies and clusters of galaxies, 13.8 billion years later, and the evolution of the expansion of space, with a relative slowdown in the matter-dominated era and, since a few billion years, an acceleration powered by dark energy. But we are far from knowing the pillars of this model which are inflation, dark matter and dark energy. Comprehending these fundamental questions requires a detailed mapping of our observable Universe over the whole of cosmic time. The relic radiation provides the starting point and galaxies draw the cosmic web. JAXA's LiteBIRD mission will map the beginning of our Universe with a crucial test for inflation (its primordial gravity waves), and the ESA Euclid mission will map the most recent half part, crucial for dark energy. The mission concept, described in this White Paper, GAUSS, aims at being a mission to fully map the cosmic web up to the reionization era, linking early and late evolution, to tackle and disentangle the crucial degeneracies persisting after the Euclid era between dark matter and inflation properties, dark energy, structure growth and gravitation at large scale.

astro-ph.CO

Fink, a new generation of broker for the LSST community

Fink is a broker designed to enable science with large time-domain alert streams such as the one from the upcoming Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST). It exhibits traditional astronomy broker features such as automatised ingestion, annotation, selection and redistribution of promising alerts for transient science. It is also designed to go beyond traditional broker features by providing real-time transient classification which is continuously improved by using state-of-the-art Deep Learning and Adaptive Learning techniques. These evolving added values will enable more accurate scientific output from LSST photometric data for diverse science cases while also leading to a higher incidence of new discoveries which shall accompany the evolution of the survey. In this paper we introduce Fink, its science motivation, architecture and current status including first science verification cases using the Zwicky Transient Facility alert stream.

astro-ph.IM

The LSST DESC DC2 Simulated Sky Survey

We describe the simulated sky survey underlying the second data challenge (DC2) carried out in preparation for analysis of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) by the LSST Dark Energy Science Collaboration (LSST DESC). Significant connections across multiple science domains will be a hallmark of LSST; the DC2 program represents a unique modeling effort that stresses this interconnectivity in a way that has not been attempted before. This effort encompasses a full end-to-end approach: starting from a large N-body simulation, through setting up LSST-like observations including realistic cadences, through image simulations, and finally processing with Rubin's LSST Science Pipelines. This last step ensures that we generate data products resembling those to be delivered by the Rubin Observatory as closely as is currently possible. The simulated DC2 sky survey covers six optical bands in a wide-fast-deep (WFD) area of approximately 300 deg^2 as well as a deep drilling field (DDF) of approximately 1 deg^2. We simulate 5 years of the planned 10-year survey. The DC2 sky survey has multiple purposes. First, the LSST DESC working groups can use the dataset to develop a range of DESC analysis pipelines to prepare for the advent of actual data. Second, it serves as a realistic testbed for the image processing software under development for LSST by the Rubin Observatory. In particular, simulated data provide a controlled way to investigate certain image-level systematic effects. Finally, the DC2 sky survey enables the exploration of new scientific ideas in both static and time-domain cosmology.

astro-ph.IM

LSST: from Science Drivers to Reference Design and Anticipated Data Products

(Abridged) We describe here the most ambitious survey currently planned in the optical, the Large Synoptic Survey Telescope (LSST). A vast array of science will be enabled by a single wide-deep-fast sky survey, and LSST will have unique survey capability in the faint time domain. The LSST design is driven by four main science themes: probing dark energy and dark matter, taking an inventory of the Solar System, exploring the transient optical sky, and mapping the Milky Way. LSST will be a wide-field ground-based system sited at Cerro Pachón in northern Chile. The telescope will have an 8.4 m (6.5 m effective) primary mirror, a 9.6 deg$^2$ field of view, and a 3.2 Gigapixel camera. The standard observing sequence will consist of pairs of 15-second exposures in a given field, with two such visits in each pointing in a given night. With these repeats, the LSST system is capable of imaging about 10,000 square degrees of sky in a single filter in three nights. The typical 5$σ$ point-source depth in a single visit in $r$ will be $\sim 24.5$ (AB). The project is in the construction phase and will begin regular survey operations by 2022. The survey area will be contained within 30,000 deg$^2$ with $δ<+34.5^\circ$, and will be imaged multiple times in six bands, $ugrizy$, covering the wavelength range 320--1050 nm. About 90\% of the observing time will be devoted to a deep-wide-fast survey mode which will uniformly observe a 18,000 deg$^2$ region about 800 times (summed over all six bands) during the anticipated 10 years of operations, and yield a coadded map to $r\sim27.5$. The remaining 10\% of the observing time will be allocated to projects such as a Very Deep and Fast time domain survey. The goal is to make LSST data products, including a relational database of about 32 trillion observations of 40 billion objects, available to the public and scientists around the world.

astro-ph

The influence of host galaxy morphology on the properties of Type Ia supernovae from the JLA compilation

The observational cosmology with distant Type Ia supernovae (SNe) as standard candles claims that the Universe is in accelerated expansion, caused by a large fraction of dark energy. In this paper we investigate the SN Ia environment, studying the impact of the nature of their host galaxies on the Hubble diagram fitting. The supernovae (192 SNe) used in the analysis were extracted from Joint-Light-curves-Analysis (JLA) compilation of high-redshift and nearby supernovae which is the best one to date. The analysis is based on the empirical fact that SN Ia luminosities depend on their light curve shapes and colors. We confirm that the stretch parameter of Type Ia supernovae is correlated with the host galaxy type. The supernovae with lower stretch are hosted mainly in elliptical and lenticular galaxies. No significant correlation between SN Ia colour and host morphology was found. We also examine how the luminosities of SNe Ia change depending on host galaxy morphology after stretch and colour corrections. Our results show that in old stellar populations and low dust environments, the supernovae are slightly fainter. SNe Ia in elliptical and lenticular galaxies have a higher $α$ (slope in luminosity-stretch) and $β$ (slope in luminosity-colour) parameter than in spirals. However, the observed shift is at the 1-$σ$ uncertainty level and, therefore, can not be considered as significant. We confirm that the supernova properties depend on their environment and that the incorporation of a host galaxy term into the Hubble diagram fit is expected to be crucial for future cosmological analyses.

astro-ph.CO