SearcharxivSearch

arXiv subjects

Stefano Cavuoti

Publications and source records attributed to Stefano Cavuoti.

At least 19 recordsLinked to original sources

Feature-driven anomaly flagging in obscured active galactic nucleus light curves with autoencoders

Active galactic nuclei (AGN) are among the most complex classes of astrophysical objects, displaying a wide range of variability and observational properties. Identifying unusual AGN is crucial for understanding the physical mechanisms behind their emission better and for discovering potentially new subclasses or rare behaviors. With the increasing volume of data from next-generation surveys, machine-learning-based anomaly detection offers a promising approach to flagging and investigating such outliers systematically. We explore the use of unsupervised algorithms with a feature-driven approach to flag anomalous AGN, further explored by a human expert. The main focus is on obscured AGN, which tend to be harder to characterize. The algorithm we used was an AutoEncoder, which we trained on features extracted from the light curves rather than working with the light curves directly. The unsupervised nature of the method allows the detection of anomalies without relying on labeled data. To properly characterize the feature space and the detection process, we used the SHAP method. Our method flagged $11.18\%$ of the AGN we studied as anomalous. We focused in particular on anomalous obscured AGN and identified a refined subset of features that yields a comparable performance to the full set. Together with an in-depth analysis of the anomalies, this provides insight into how the AutoEncoder assigns anomalous status and which features are most indicative of astrophysically interesting behaviors or phenomena.

astro-ph.GA

Classification of blazars based on data-driven approaches

Active galactic nuclei (AGNs), including blazars, exhibit distinctive variability in their optical light curves, making them ideal for classification studies. This work uses data from the latest GAIA and Pan-STARRS data releases to analyze these patterns. The goal of this work is to classify AGNs into two categories: "blazars" and "non-blazars'' using only optical light curves. This strategy differs from most existing works, as it relies exclusively on optical variability without employing any other multiwavelength information. We processed optical light curves from GAIA and Pan-STARRS using the FATS library to extract standard time-series features. We computed additional features with custom algorithms based on literature methods. A Light Gradient-Boosting Machine (LightGBM) model was trained to classify AGNs into blazars and non-blazars based on these features. We then used this knowledge base to carry out a self-learning experiment with AGN candidates of an unknown nature. The LightGBM model achieved an accuracy of $86\%$, with precision, recall, and F1 score above $80-85\%$ for classifying blazars and non-blazar AGNs using optical data. The application of a BoostBoruta algorithm for feature selection reduced the feature space from 70 to 13. while maintaining comparable performance. A self-training classifier yielded similar results $85\%$, confirming the robustness of the model and the reliability of pseudo-labeling for unknown objects.

astro-ph.GA

A Quantum Genetic Algorithm with application to Cosmological Parameters Estimation

An Amplitude-Encoded Quantum Genetic Algorithm (AEQGA) has been developed to minimize $χ^2$ functions of different cosmological probes (Supernovae Type Ia, Baryon Acoustic Oscillations, Cosmic Microwave Background Radiation), to find the best-fit value for two cosmological parameters, namely the Hubble Constant and the density matter content of the Universe today. Our main aim is to pave the way to testing the adoption of quantum optimization in the inference of the cosmological parameters that describe the universe evolution. AEQGA computes the merit function classically, and then uses a quantum circuit to entangle the population and perform crossover and mutation operations. The results show consistency with the isocontours of the objective functions. We then tested the general behavior of AEQGA as a function of its hyperparameters and compared it with a second quantum genetic algorithm found in the literature as well as with classical algorithms, finding consistent results.

astro-ph.CO

Scavenger hunt: Selection of obscured active galactic nuclei combining multiband optical variability and colors

As wide-field optical surveys such as Vera Rubin Observatory's Legacy Survey of Space and Time (LSST) begin operations, time-domain astronomy is facing a data revolution, paving the road for new, expanded variability studies. This work leverages the complementary power of optical variability and color selection to identify active galactic nuclei (AGN), focusing on optimizing the identification of obscured AGN, typically more challenging to distinguish from inactive galaxies based on optical variability alone. The analysis is designed to provide valuable insights in the context of performance preview for the LSST, albeit using a scaled-down version of the LSST dataset. We present the first combined AGN selection based on g+r+i band light curves from the VST-COSMOS survey, spanning 3.3 yr. We identify AGN candidates independently in each band using a random forest (RF) classifier trained on features mainly related to optical variability, along with six optical/infrared colors and a morphology indicator. We subsequently merge the three band-specific samples in order to enhance selection purity and reliability. We then focus on defining a subset of features that significantly improve the identification of obscured AGN. The RF classifiers yield a consistent performance across the three bands, highlighting the critical role of contamination. Using the combined three-band plus color selection we successfully recover $58^{+9}_{-8}\%$ of all AGN and $69^{+10}_{-8}\%$ of the known obscured AGN that have been independently confirmed in all three bands. When requiring confirmation in two out of the three bands, these fractions increase to $69^{+10}_{-8}\%$ and $80^{+10}_{-9}\%$, respectively. We also demonstrate that, while combining variability features with colors is crucial to improve obscured AGN selection, relying solely on color features returns a markedly higher contamination rate.

astro-ph.GA

Quantum Markov Chain Monte Carlo for Cosmological Functions

We present an implementation of Quantum Computing for a Markov Chain Monte Carlo method with an application to cosmological functions, to derive posterior distributions from cosmological probes. The algorithm proposes new steps in the parameter space via a quantum circuit whose resulting statevector provides the components of the shift vector. The proposed point is accepted or rejected via the classical Metropolis-Hastings acceptance method. The advantage of this hybrid quantum approach is that the step size and direction change in a way independent of the evolution of the chain, thus ideally avoiding the presence of local minima. The results are consistent with analyses performed with classical methods, both for a test function and real cosmological data. The final goal is to generalize this algorithm to test its application to complex cosmological computations.

astro-ph.CO

ULISSE: Determination of star-formation rate and stellar mass based on the one-shot galaxy imaging technique

Modern sky surveys produce vast amounts of observational data, making the application of classical methods for estimating galaxy properties challenging and time-consuming. This challenge can be significantly alleviated by employing automatic machine and deep learning techniques. We propose an implementation of the ULISSE algorithm aimed at determining physical parameters of galaxies, in particular star-formation rates (SFR) and stellar masses ($M_{\ast}$), using only composite-color images. ULISSE is able to rapidly and efficiently identify candidates from a single image based on photometric and morphological similarities to a given reference object with known properties. This approach leverages features extracted from the ImageNet dataset to perform similarity searches among all objects in the sample, eliminating the need for extensive neural network training. Our experiments, performed on the Sloan Digital Sky Survey, demonstrate that we are able to predict the joint star formation rate and stellar mass of the target galaxies within 1 dex in 60% to 80% of cases, depending on the investigated subsample (quiescent/star-forming galaxies, early-/late-type, etc.), and within 0.5 dex if we consider these parameters separately. This is approximately twice the fraction obtained from a random guess extracted from the parent population. Additionally, we find ULISSE is more effective for galaxies with active star formation compared to elliptical galaxies with quenched star formation. Additionally, ULISSE performs more efficiently for galaxies with bright nuclei such as AGN. Our results suggest that ULISSE is a promising tool for a preliminary estimation of star-formation rates and stellar masses for galaxies based only on single images in current and future wide-field surveys (e.g., Euclid, LSST), which target millions of sources nightly.

astro-ph.GA

Navigating AGN variability with self-organizing maps

Context. The classification of active galactic nuclei (AGNs) is a challenge in astrophysics. Variability features extracted from light curves offer a promising avenue for distinguishing AGNs and their subclasses. This approach would be very valuable in sight of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST). Aims. Our goal is to utilize self-organizing maps (SOMs) to classify AGNs based on variability features and investigate how the use of different subsets of features impacts the purity and completeness of the resulting classifications. Methods. We derived a set of variability features from light curves, similar to those employed in previous studies, and applied SOMs to explore the distribution of AGNs subclasses. We conducted a comparative analysis of the classifications obtained with different subsets of features, focusing on the ability to identify different AGNs types. Results. Our analysis demonstrates that using SOMs with variability features yields a relatively pure AGNs sample, though completeness remains a challenge. In particular, Type 2 AGNs are the hardest to identify, as can be expected. These results represent a promising step toward the development of tools that may support AGNs selection in future large-scale surveys such as LSST.

astro-ph.IM

The Evolutionary Path of Star-Forming Clumps in Hi-GAL

Star formation (SF) studies are benefiting from the huge amount of data made available by recent large-area Galactic plane surveys conducted between 2 μm and 3 mm. Fully characterizing SF demands integrating far-infrared/sub-millimetre (FIR/sub-mm) data, tracing the earliest phases, with near-/mid-infrared (NIR/MIR) observations, revealing later stages characterized by YSOs just before main sequence star appearance. However, the resulting dataset is often a complex mix of heterogeneous and intricate features, limiting the effectiveness of traditional analysis in uncovering hidden patterns and relationships. In this framework, machine learning emerges as a powerful tool to handle the complexity of feature-rich datasets and investigate potential physical connections between the cold dust component traced by FIR/sub-mm emission and the presence of YSOs. We present a study on the evolutionary path of star forming clumps in the Hi-GAL survey through a multi-step approach, with the final aims of (a) obtaining a robust and accurate set of features able to well classify the star forming clumps in Hi-GAL based on their evolutionary properties, (b) establishing whether a connection exists between the cold material reservoir in clumps, traced by FIR/sub-mm emission, and the already formed YSOs, precursors of stars. For these purposes, our designed experiments aim at testing whether the FIR/sub-mm properties related to clumps are sufficient to predict the clump evolutionary stage, without considering the direct information about the embedded YSOs at NIR/MIR. Our machine learning-based method involves a four-step approach, based on feature engineering, data handling, feature selection and classification. Our findings suggest that FIR/sub-mm and NIR/MIR emissions trace different evolutionary phases of star forming clumps, highlighting the complex and asynchronous nature of the SF process.

astro-ph.GA

Selection of optically variable active galactic nuclei via a random forest algorithm

Context. A defining characteristic of active galactic nuclei (AGN) that distinguishes them from other astronomical sources is their stochastic variability, which is observable across the entire electromagnetic spectrum. Upcoming optical wide-field surveys, such as the Vera C. Rubin Observatory's Legacy Survey of Space and Time, are set to transform astronomy by delivering unprecedented volumes of data for time domain studies. This data influx will require the development of the expertise and methodologies necessary to manage and analyze it effectively. Aims. This project focuses on optimizing AGN selection through optical variability in wide-field surveys and aims to reduce the bias against obscured AGN. We tested a random forest (RF) algorithm trained on various feature sets to select AGN. The initial dataset consisted of 54 observations in the r-band and 25 in the g-band of the COSMOS field, captured with the VLT Survey Telescope over a 3.3-year baseline. Methods. Our analysis relies on feature sets derived separately from either band plus a set of features combining data from both bands, mostly characterizing AGN on the basis of their variability properties and obtained from their light curves. We trained multiple RF classifiers using different subsets of selected features and assessed their performance via targeted metrics. Results. Our tests provide valuable insights into the use of multiband and multivisit data for AGN identification. We compared our findings with previous studies and dedicated part of the analysis to potential enhancements in selecting obscured AGN. The expertise gained and the methodologies developed here are readily applicable to datasets from other ground- and space-based missions.

astro-ph.GA

AMBER -- Advanced SegFormer for Multi-Band Image Segmentation: an application to Hyperspectral Imaging

Deep learning has revolutionized the field of hyperspectral image (HSI) analysis, enabling the extraction of complex spectral and spatial features. While convolutional neural networks (CNNs) have been the backbone of HSI classification, their limitations in capturing global contextual features have led to the exploration of Vision Transformers (ViTs). This paper introduces AMBER, an advanced SegFormer specifically designed for multi-band image segmentation. AMBER enhances the original SegFormer by incorporating three-dimensional convolutions, custom kernel sizes, and a Funnelizer layer. This architecture enables processing hyperspectral data directly, without requiring spectral dimensionality reduction during preprocessing. Our experiments, conducted on three benchmark datasets (Salinas, Indian Pines, and Pavia University) and on a dataset from the PRISMA satellite, show that AMBER outperforms traditional CNN-based methods in terms of Overall Accuracy, Kappa coefficient, and Average Accuracy on the first three datasets, and achieves state-of-the-art performance on the PRISMA dataset. These findings highlight AMBER's robustness, adaptability to both airborne and spaceborne data, and its potential as a powerful solution for remote sensing and other domains requiring advanced analysis of high-dimensional data.

cs.CV

The fifth data release of the Kilo Degree Survey: Multi-epoch optical/NIR imaging covering wide and legacy-calibration fields

We present the final data release of the Kilo-Degree Survey (KiDS-DR5), a public European Southern Observatory (ESO) wide-field imaging survey optimised for weak gravitational lensing studies. We combined matched-depth multi-wavelength observations from the VLT Survey Telescope and the VISTA Kilo-degree INfrared Galaxy (VIKING) survey to create a nine-band optical-to-near-infrared survey spanning $1347$ deg$^2$. The median $r$-band $5σ$ limiting magnitude is 24.8 with median seeing $0.7^{\prime\prime}$. The main survey footprint includes $4$ deg$^2$ of overlap with existing deep spectroscopic surveys. We complemented these data in DR5 with a targeted campaign to secure an additional $23$ deg$^2$ of KiDS- and VIKING-like imaging over a range of additional deep spectroscopic survey fields. From these fields, we extracted a catalogue of $126\,085$ sources with both spectroscopic and photometric redshift information, which enables the robust calibration of photometric redshifts across the full survey footprint. In comparison to previous releases, DR5 represents a $34\%$ areal extension and includes an $i$-band re-observation of the full footprint, thereby increasing the effective $i$-band depth by $0.4$ magnitudes and enabling multi-epoch science. Our processed nine-band imaging, single- and multi-band catalogues with masks, and homogenised photometry and photometric redshifts can be accessed through the ESO Archive Science Portal.

astro-ph.GA

Leveraging Transfer Learning for Astronomical Image Analysis

The exponential growth of astronomical data from large-scale surveys has created both opportunities and challenges for the astrophysics community. This paper explores the possibilities offered by transfer learning techniques in addressing these challenges across various domains of astronomical research. We present a set of recent applications of transfer learning methods for astronomical tasks based on the usage of a pre-trained convolutional neural networks. The examples shortly discussed include the detection of candidate active galactic nuclei (AGN), the possibility of deriving physical parameters for galaxies directly from images, the identification of artifacts in time series images, and the detection of strong lensing candidates and outliers. We demonstrate how transfer learning enables efficient analysis of complex astronomical phenomena, particularly in scenarios where labeled data is scarce. This kind of method will be very helpful for upcoming large-scale surveys like the Rubin Legacy Survey of Space and Time (LSST). By showcasing successful implementations and discussing methodological approaches, we highlight the versatility and effectiveness of such techniques.

astro-ph.IM

Galaxy spectroscopy without spectra: Galaxy properties from photometric images with conditional diffusion models

Modern spectroscopic surveys can only target a small fraction of the vast amount of photometrically cataloged sources in wide-field surveys. Here, we report the development of a generative AI method capable of predicting optical galaxy spectra from photometric broad-band images alone. This method draws from the latest advances in diffusion models in combination with contrastive networks. We pass multi-band galaxy images into the architecture to obtain optical spectra. From these, robust values for galaxy properties can be derived with any methods in the spectroscopic toolbox, such as standard population synthesis techniques and Lick indices. When trained and tested on 64x64-pixel images from the Sloan Digital Sky Survey, the global bimodality of star-forming and quiescent galaxies in photometric space is recovered, as well as a mass-metallicity relation of star-forming galaxies. The comparison between the observed and the artificially created spectra shows good agreement in overall metallicity, age, Dn4000, stellar velocity dispersion, and E(B-V) values. Photometric redshift estimates of our generative algorithm can compete with other current, specialized deep-learning techniques. Moreover, this work is the first attempt in the literature to infer velocity dispersion from photometric images. Additionally, we can predict the presence of an active galactic nucleus up to an accuracy of 82%. With our method, scientifically interesting galaxy properties, normally requiring spectroscopic inputs, can be obtained in future data sets from large-scale photometric surveys alone. The spectra prediction via AI can further assist in creating realistic mock catalogs.

astro-ph.GA

Identification of problematic epochs in astronomical time series through transfer learning

We present a novel method for detecting outliers in astronomical time series based on the combination of a deep neural network and a k-nearest neighbor algorithm with the aim of identifying and removing problematic epochs in the light curves of astronomical objects. We use an EfficientNet network pre-trained on ImageNet as a feature extractor and perform a k-nearest neighbor search in the resulting feature space to measure the distance from the first neighbor for each image. If the distance is above the one obtained for a stacked image, we flag the image as a potential outlier. We apply our method to time series obtained from the VLT Survey Telescope (VST) monitoring campaign of the Deep Drilling Fields of the Vera C. Rubin Legacy Survey of Space and Time (LSST). We show that our method can effectively identify and remove artifacts from the VST time series and improve the quality and reliability of the data. This approach may prove very useful in sight of the amount of data that will be provided by the LSST, which will prevent the inspection of individual light curves. We also discuss the advantages and limitations of our method and suggest possible directions for future work.

astro-ph.IM

Data-Driven Approaches to Searches for the Technosignatures of Advanced Civilizations

Humanity has wondered whether we are alone for millennia. The discovery of life elsewhere in the Universe, particularly intelligent life, would have profound effects, comparable to those of recognizing that the Earth is not the center of the Universe and that humans evolved from previous species. There has been rapid growth in the fields of extrasolar planets and data-driven astronomy. In a relatively short interval, we have seen a change from knowing of no extrasolar planets to now knowing more potentially habitable extrasolar planets than there are planets in the Solar System. In approximately the same interval, astronomy has transitioned to a field in which sky surveys can generate 1 PB or more of data. The Data-Driven Approaches to Searches for the Technosignatures of Advanced Civilizations_ study at the W. M. Keck Institute for Space Studies was intended to revisit searches for evidence of alien technologies in light of these developments. Data-driven searches, being able to process volumes of data much greater than a human could, and in a reproducible manner, can identify *anomalies* that could be clues to the presence of technosignatures. A key outcome of this workshop was that technosignature searches should be conducted in a manner consistent with Freeman Dyson's "First Law of SETI Investigations," namely "every search for alien civilizations should be planned to give interesting results even when no aliens are discovered." This approach to technosignatures is commensurate with NASA's approach to biosignatures in that no single observation or measurement can be taken as providing full certainty for the detection of life. Areas of particular promise identified during the workshop were (*) Data Mining of Large Sky Surveys, (*) All-Sky Survey at Far-Infrared Wavelengths, (*) Surveys with Radio Astronomical Interferometers, and (*) Artifacts in the Solar System.

astro-ph.IM

Generating astronomical spectra from photometry with conditional diffusion models

A trade-off between speed and information controls our understanding of astronomical objects. Fast-to-acquire photometric observations provide global properties, while costly and time-consuming spectroscopic measurements enable a better understanding of the physics governing their evolution. Here, we tackle this problem by generating spectra directly from photometry, through which we obtain an estimate of their intricacies from easily acquired images. This is done by using multi-modal conditional diffusion models, where the best out of the generated spectra is selected with a contrastive network. Initial experiments on minimally processed SDSS galaxy data show promising results.

astro-ph.IM

ULISSE: A Tool for One-shot Sky Exploration and its Application to Active Galactic Nuclei Detection

Modern sky surveys are producing ever larger amounts of observational data, which makes the application of classical approaches for the classification and analysis of objects challenging and time-consuming. However, this issue may be significantly mitigated by the application of automatic machine and deep learning methods. We propose ULISSE, a new deep learning tool that, starting from a single prototype object, is capable of identifying objects sharing the same morphological and photometric properties, and hence of creating a list of candidate sosia. In this work, we focus on applying our method to the detection of AGN candidates in a Sloan Digital Sky Survey galaxy sample, since the identification and classification of Active Galactic Nuclei (AGN) in the optical band still remains a challenging task in extragalactic astronomy. Intended for the initial exploration of large sky surveys, ULISSE directly uses features extracted from the ImageNet dataset to perform a similarity search. The method is capable of rapidly identifying a list of candidates, starting from only a single image of a given prototype, without the need for any time-consuming neural network training. Our experiments show ULISSE is able to identify AGN candidates based on a combination of host galaxy morphology, color and the presence of a central nuclear source, with a retrieval efficiency ranging from 21% to 65% (including composite sources) depending on the prototype, where the random guess baseline is 12%. We find ULISSE to be most effective in retrieving AGN in early-type host galaxies, as opposed to prototypes with spiral- or late-type properties. Based on the results described in this work, ULISSE can be a promising tool for selecting different types of astrophysical objects in current and future wide-field surveys (e.g. Euclid, LSST etc.) that target millions of sources every single night.

astro-ph.IM

Galaxy morphoto-Z with neural Networks (GaZNets). I. Optimized accuracy and outlier fraction from Imaging and Photometry

In the era of large sky surveys, photometric redshifts (photo-z) represent crucial information for galaxy evolution and cosmology studies. In this work, we propose a new Machine Learning (ML) tool called Galaxy morphoto-Z with neural Networks (GaZNet-1), which uses both images and multi-band photometry measurements to predict galaxy redshifts, with accuracy, precision and outlier fraction superior to standard methods based on photometry only. As a first application of this tool, we estimate photo-z of a sample of galaxies in the Kilo-Degree Survey (KiDS). GaZNet-1 is trained and tested on $\sim140 000$ galaxies collected from KiDS Data Release 4 (DR4), for which spectroscopic redshifts are available from different surveys. This sample is dominated by bright (MAG$\_$AUTO$<21$) and low redshift ($z < 0.8$) systems, however, we could use $\sim$ 6500 galaxies in the range $0.8 < z < 3$ to effectively extend the training to higher redshift. The inputs are the r-band galaxy images plus the 9-band magnitudes and colours, from the combined catalogs of optical photometry from KiDS and near-infrared photometry from the VISTA Kilo-degree Infrared survey. By combining the images and catalogs, GaZNet-1 can achieve extremely high precision in normalized median absolute deviation (NMAD=0.014 for lower redshift and NMAD=0.041 for higher redshift galaxies) and low fraction of outliers ($0.4$\% for lower and $1.27$\% for higher redshift galaxies). Compared to ML codes using only photometry as input, GaZNet-1 also shows a $\sim 10-35$% improvement in precision at different redshifts and a $\sim$ 45% reduction in the fraction of outliers. We finally discuss that, by correctly separating galaxies from stars and active galactic nuclei, the overall photo-z outlier fraction of galaxies can be cut down to $0.3$\%.

astro-ph.GA