SearcharxivSearch

arXiv subjects

Sreevarsha Sreejith

Publications and source records attributed to Sreevarsha Sreejith.

15 recordsLinked to original sources

SNAD: enabling discovery in the era of big data

In the era of wide-field surveys and big data in astronomy, the SNAD team is exploiting the potential of modern datasets for discovering new, unforeseen, or rare astrophysical objects and phenomena with machine learning (ML). The SNAD pipeline was built under the hypothesis that, although automatic ML algorithms have a crucial role to play in this task, the scientific discovery is only completely realized when such systems are designed to boost the impact of domain knowledge experts. Our key contributions include the development of the Coniferest Python library, which offers implementations of two active learning algorithms with an ``expert in loop'', and the creation of the SNAD Transient Miner, facilitating the search for specific types of transients. We have also developed the SNAD Viewer, a web portal that provides a centralized view of individual objects from the Zwicky Transient Facility's (ZTF) data releases, making the analysis of potential anomalies more efficient. Finally, when applied to ZTF data, our approach has resulted in more than a hundred new supernova (SN) candidates, along with a few other non-catalogued objects, such as red dwarf flares, superluminous SNe, RS CVn type variables, and young stellar objects.

astro-ph.HE

What ZTF Saw Where Rubin Looked: Anomaly Hunting in DR23

We present results from the SNAD VIII Workshop, during which we conducted the first systematic anomaly search in the ZTF fields also observed by LSSTComCam during Rubin Scientific Pipeline commissioning. Using the PineForest active anomaly detection algorithm, we analysed four selected fields (two galactic and two extragalactic) and visually inspected 400 candidates. As a result, we discovered six previously uncatalogued variable stars, including RS~CVn, BY Draconis, ellipsoidal, and solar-type variables, and refined classifications and periods for six known objects. These results demonstrate the effectiveness of the SNAD anomaly detection pipeline and provide a preview of the discovery potential in the upcoming LSST data.

astro-ph.IM

Signatures to help interpretability of anomalies

Machine learning is often viewed as a black box when it comes to understanding its output, be it a decision or a score. Automatic anomaly detection is no exception to this rule, and quite often the astronomer is left to independently analyze the data in order to understand why a given event is tagged as an anomaly. We introduce here idea of anomaly signature, whose aim is to help the interpretability of anomalies by highlighting which features contributed to the decision.

cs.LG

Dataset of artefacts for machine learning applications in astronomy

Accurate photometry in astronomical surveys is challenged by image artefacts, which affect measurements and degrade data quality. Due to the large amount of available data, this task is increasingly handled using machine learning algorithms, which often require a labelled training set to learn data patterns. We present an expert-labelled dataset of 1127 artefacts with 1213 labels from 26 fields in ZTF DR3, along with a complementary set of nominal objects. The artefact dataset was compiled using the active anomaly detection algorithm PineForest, developed by the SNAD team. These datasets can serve as valuable resources for real-bogus classification, catalogue cleaning, anomaly detection, and educational purposes. Both artefacts and nominal images are provided in FITS format in two sizes (28 x 28 and 63 x 63 pixels). The datasets are publicly available for further scientific applications.

astro-ph.IM

Exploring the Universe with SNAD: Anomaly Detection in Astronomy

SNAD is an international project with a primary focus on detecting astronomical anomalies within large-scale surveys, using active learning and other machine learning algorithms. The work carried out by SNAD not only contributes to the discovery and classification of various astronomical phenomena but also enhances our understanding and implementation of machine learning techniques within the field of astrophysics. This paper provides a review of the SNAD project and summarizes the advancements and achievements made by the team over several years.

astro-ph.IM

Point Spread Function Deconvolution Using a Convolutional Autoencoder for Astronomical Applications

A major issue in optical astronomical image analysis is the combined effect of the instrument's point spread function (PSF) and the atmospheric seeing that blurs images and changes their shape in a way that is band and time-of-observation dependent. In this work we present a very simple neural network based approach to non-blind image deconvolution that relies on feeding a Convolutional Autoencoder (CAE) input images that have been preprocessed by convolution with the corresponding PSF and its regularized inverse, a method which is both conceptually simple and computationally less intensive. We also present here, a new approach for dealing with limited input dynamic range of neural networks compared to the dynamic range present in astronomical images.

astro-ph.IM

Neural Network Based Point Spread Function Deconvolution For Astronomical Applications

Optical astronomical images are strongly affected by the point spread function (PSF) of the optical system and the atmosphere (seeing) which blurs the observed image. The amount of blurring depends both on the observed band, and on the atmospheric conditions during observation. A typical astronomical image will likely have a unique PSF, that is non-circular and different in different bands. At the same time, observations of known stars also give us an accurate determination of this PSF. Therefore, any serious candidate for production analysis of astronomical images must take the known PSF into account during the image analysis. So far, the majority of applications of neural networks (NN) to astronomical image analysis have ignored this problem by assuming a fixed PSF in training and validation. We present a neural-network based deconvolution algorithm based on Deep Wiener Deconvolution Network (DWDN). This algorithm belongs to a class of non-blind deconvolution algorithms, since it assumes the PSF shape is known. We study the performance of different versions of this algorithm under realistic observational conditions in terms of the recovery of the most relevant astronomical quantities such as colors, ellipticities and orientations. We investigate custom loss functions that optimize the recovery of astronomical quantities with mixed results.

astro-ph.IM

Are classification metrics good proxies for SN Ia cosmological constraining power?

Context: When selecting a classifier to use for a supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, i.e. contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would save those designing an analysis pipeline from the computational expense of a full cosmology forecast. Aims: This study tests the assumption that classification metrics are an appropriate proxy for cosmology metrics. Methods: We emulate photometric SN Ia cosmology samples with controlled contamination rates of individual contaminant classes and evaluate each of them under a set of classification metrics. We then derive cosmological parameter constraints from all samples under two common analysis approaches and quantify the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results: We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are insensitive to the latter. Conclusions: We therefore discourage exclusive reliance on classification-based metrics for cosmological analysis design decisions, e.g. classifier choice, and instead recommend optimizing using a metric of cosmological parameter constraining power.

astro-ph.CO

Supernova search with active learning in ZTF DR3

We provide the first results from the complete SNAD adaptive learning pipeline in the context of a broad scope of data from large-scale astronomical surveys. The main goal of this work is to explore the potential of adaptive learning techniques in application to big data sets. Our SNAD team used Active Anomaly Discovery (AAD) as a tool to search for new supernova (SN) candidates in the photometric data from the first 9.4 months of the Zwicky Transient Facility (ZTF) survey, namely, between March 17 and December 31 2018 (58194 < MJD < 58483). We analysed 70 ZTF fields at a high galactic latitude and visually inspected 2100 outliers. This resulted in 104 SN-like objects being found, 57 of which were reported to the Transient Name Server for the first time and with 47 having previously been mentioned in other catalogues, either as SNe with known types or as SN candidates. We visually inspected the multi-colour light curves of the non-catalogued transients and performed fittings with different supernova models to assign it to a probable photometric class: Ia, Ib/c, IIP, IIL, or IIn. Moreover, we also identified unreported slow-evolving transients that are good superluminous SN candidates, along with a few other non-catalogued objects, such as red dwarf flares and active galactic nuclei. Beyond confirming the effectiveness of human-machine integration underlying the AAD strategy, our results shed light on potential leaks in currently available pipelines. These findings can help avoid similar losses in future large-scale astronomical surveys. Furthermore, the algorithm enables direct searches of any type of data and based on any definition of an anomaly set by the expert.

astro-ph.HE

The SNAD Viewer: Everything You Want to Know about Your Favorite ZTF Object

We describe the SNAD Viewer, a web portal for astronomers which presents a centralized view of individual objects from the Zwicky Transient Facility's (ZTF) data releases, including data gathered from multiple publicly available astronomical archives and data sources. Initially built to enable efficient expert feedback in the context of adaptive machine learning applications, it has evolved into a full-fledged community asset that centralizes public information and provides a multi-dimensional view of ZTF sources. For users, we provide detailed descriptions of the data sources and choices underlying the information displayed in the portal. For developers, we describe our architectural choices and their consequences such that our experience can help others engaged in similar endeavors or in adapting our publicly released code to their requirements. The infrastructure we describe here is scalable and flexible and can be personalized and used by other surveys and for other science goals. The Viewer has been instrumental in highlighting the crucial roles domain experts retain in the era of big data in astronomy. Given the arrival of the upcoming generation of large-scale surveys, we believe similar systems will be paramount in enabling an optimal exploitation of the scientific potential enclosed in current terabyte and future petabyte-scale data sets. The Viewer is publicly available online at https://ztf.snad.space

astro-ph.IM

Galaxy Deblending using Residual Dense Neural networks

We present a new neural network approach for deblending galaxy images in astronomical data using Residual Dense Neural network (RDN) architecture. We train the network on synthetic galaxy images similar to the typical arrangements of field galaxies with a finite point spread function (PSF) and realistic noise levels. The main novelty of our approach is the usage of two distinct neural networks: i) a deblending network which isolates a single galaxy postage stamp from the composite and, ii) a classifier network which counts the remaining number of galaxies. The deblending proceeds by iteratively peeling one galaxy at a time from the composite until the image contains no further objects as determined by the classifier, or by other stopping criteria. By looking at the consistency in the outputs of the two networks, we can assess the quality of the deblending. We characterize the flux and shape reconstructions in different quality bins and compare our deblender with the industry standard, SExtractor. We also discuss possible future extensions for the project with variable PSFs and noise levels.

astro-ph.GA

Active learning with RESSPECT: Resource allocation for extragalactic astronomical transients

The recent increase in volume and complexity of available astronomical data has led to a wide use of supervised machine learning techniques. Active learning strategies have been proposed as an alternative to optimize the distribution of scarce labeling resources. However, due to the specific conditions in which labels can be acquired, fundamental assumptions, such as sample representativeness and labeling cost stability cannot be fulfilled. The Recommendation System for Spectroscopic follow-up (RESSPECT) project aims to enable the construction of optimized training samples for the Rubin Observatory Legacy Survey of Space and Time (LSST), taking into account a realistic description of the astronomical data environment. In this work, we test the robustness of active learning techniques in a realistic simulated astronomical data scenario. Our experiment takes into account the evolution of training and pool samples, different costs per object, and two different sources of budget. Results show that traditional active learning strategies significantly outperform random sampling. Nevertheless, more complex batch strategies are not able to significantly overcome simple uncertainty sampling techniques. Our findings illustrate three important points: 1) active learning strategies are a powerful tool to optimize the label-acquisition task in astronomy, 2) for upcoming large surveys like LSST, such techniques allow us to tailor the construction of the training sample for the first day of the survey, and 3) the peculiar data environment related to the detection of astronomical transients is a fertile ground that calls for the development of tailored machine learning algorithms.

astro-ph.IM

Active Anomaly Detection for time-domain discoveries

We present the first evidence that adaptive learning techniques can boost the discovery of unusual objects within astronomical light curve data sets. Our method follows an active learning strategy where the learning algorithm chooses objects which can potentially improve the learner if additional information about them is provided. This new information is subsequently used to update the machine learning model, allowing its accuracy to evolve with each new information. For the case of anomaly detection, the algorithm aims to maximize the number of scientifically interesting anomalies presented to the expert by slightly modifying the weights of a traditional Isolation Forest (IF) at each iteration. In order to demonstrate the potential of such techniques, we apply the Active Anomaly Discovery (AAD) algorithm to 2 data sets: simulated light curves from the PLAsTiCC challenge and real light curves from the Open Supernova Catalog. We compare the AAD results to those of a static IF. For both methods, we performed a detailed analysis for all objects with the ~2% highest anomaly scores. We show that, in the real data scenario, AAD was able to identify ~80\% more true anomalies than the IF. This result is the first evidence that AAD algorithms can play a central role in the search for new physics in the era of large scale sky surveys.

astro-ph.IM

Newly discovered dwarf galaxies in the MATLAS low density fields

We present the photometric properties of 2210 newly identified dwarf galaxy candidates in the MATLAS fields. The Mass Assembly of early Type gaLAxies with their fine Structures (MATLAS) deep imaging survey mapped $\sim$142 deg$^2$ of the sky around nearby isolated early type galaxies using MegaCam on the Canada-France-Hawaii Telescope, reaching surface brightnesses of $\sim$ 28.5 - 29 in the g-band. The dwarf candidates were identified through a direct visual inspection of the images and by visually cleaning a sample selected using a partially automated approach, and were morphologically classified at the time of identification. Approximately 75% of our candidates are dEs, indicating that a large number of early type dwarfs also populate low density environments, and 23.2% are nucleated. Distances were determined for 13.5% of our sample using pre-existing $z_{spec}$ measurements and HI detections. We confirm the dwarf nature for 99% of this sub-sample based on a magnitude cut $M_g$ = -18. Additionally, most of these ($\sim$90%) have relative velocities suggesting that they form a satellite population around nearby massive galaxies rather than an isolated field sample. Assuming that the candidates over the whole survey are satellites of the nearby galaxies, we demonstrate that the MATLAS dwarfs follow the same scaling relations as dwarfs in the Local Group as well as the Virgo and Fornax clusters. We also find that the nucleated fraction increases with $M_g$, and find evidence of a morphology-density relation for dwarfs around isolated massive galaxies.

astro-ph.GA

Galaxy And Mass Assembly: Automatic Morphological Classification of Galaxies Using Statistical Learning

We apply four statistical learning methods to a sample of $7941$ galaxies ($z<0.06$) from the Galaxy and Mass Assembly (GAMA) survey to test the feasibility of using automated algorithms to classify galaxies. Using $10$ features measured for each galaxy (sizes, colours, shape parameters \& stellar mass) we apply the techniques of Support Vector Machines (SVM), Classification Trees (CT), Classification Trees with Random Forest (CTRF) and Neural Networks (NN), returning True Prediction Ratios (TPRs) of $75.8\%$, $69.0\%$, $76.2\%$ and $76.0\%$ respectively. Those occasions whereby all four algorithms agree with each other yet disagree with the visual classification (`unanimous disagreement') serves as a potential indicator of human error in classification, occurring in $\sim9\%$ of ellipticals, $\sim9\%$ of Little Blue Spheroids, $\sim14\%$ of early-type spirals, $\sim21\%$ of intermediate-type spirals and $\sim4\%$ of late-type spirals \& irregulars. We observe that the choice of parameters rather than that of algorithms is more crucial in determining classification accuracy. Due to its simplicity in formulation and implementation, we recommend the CTRF algorithm for classifying future galaxy datasets. Adopting the CTRF algorithm, the TPRs of the 5 galaxy types are : E, $70.1\%$; LBS, $75.6\%$; S0-Sa, $63.6\%$; Sab-Scd, $56.4\%$ and Sd-Irr, $88.9\%$. Further, we train a binary classifier using this CTRF algorithm that divides galaxies into spheroid-dominated (E, LBS \& S0-Sa) and disk-dominated (Sab-Scd \& Sd-Irr), achieving an overall accuracy of $89.8\%$. This translates into an accuracy of $84.9\%$ for spheroid-dominated systems and $92.5\%$ for disk-dominated systems.

astro-ph.GA