SearcharxivSearch

arXiv subjects

Verlon Etsebeth

Publications and source records attributed to Verlon Etsebeth.

6 recordsLinked to original sources

The promise of self-supervised and active learning for Strong Lens discovery: Astronomaly applied to KiDS

Strong gravitational lenses (SGLs) are rare systems whose discovery currently relies primarily on supervised machine learning methods trained on large simulated datasets. We present the first application of Astronomaly:PROTEGE to SGL discovery, demonstrating that a human-in-the-loop active learning framework can efficiently identify lenses in large imaging surveys without the need for simulated training data. We consider a sample of 3.7 million bright galaxies from the Kilo-Degree Survey (KiDS) DR4. Feature representations are extracted using a convolutional neural network pre-trained on the ImageNet dataset and subsequently fine-tuned on KiDS data using the self-supervised Bootstrap Your Own Latent (BYOL) framework. Within the embedding of these representations, the active learning loop of Astronomaly iteratively selects the most informative systems for expert inspection. A total of 3,000 objects are inspected across multiple rounds, yielding 34 high-quality (grade A/B) SGL candidates. On the basis that these systems occupy similar regions in the learned feature space, we expand this sample through nearest-neighbour similarity analysis. Including the active learning discoveries, we identify a total of 140 grade A/B candidates and more than 1,000 additional lower-confidence systems (grade C). Among the A/B candidates, 81 are newly identified, while approximately 22% of previously known grade A/B KiDS lenses are recovered. These results demonstrate strong potential for next-generation surveys such as Euclid, Roman, and Rubin's Legacy Survey of Space and Time. With approximately 60% of the high-quality candidates newly reported, this approach complements supervised methods by reducing reliance on simulations and enabling the discovery of a diverse population of SGLs.

astro-ph.IM

A targeted machine learning approach for detecting diffuse radio emission with Astronomaly: Protege

Diffuse radio emission in galaxy clusters, such as radio halos, relics, and mini halos, is a key tracer of non-thermal processes, turbulence, and magnetic fields within the intra-cluster medium. However, their low surface brightness, as well as contamination from compact sources and imaging artefacts, makes their detection challenging. The sheer volume of data from instruments such as the Square Kilometre Array will render traditional manual-inspection based detection methods infeasible. This paper introduces a novel machine learning approach that uses active learning to rapidly identify diffuse emission candidates from a small, optimally-selected subset of data. We apply the self-supervised deep learning algorithm Bootstrap Your Own Latent to extract features from source cutouts in the MeerKAT Galaxy Cluster Legacy Survey (MGCLS). We then pass these features through the Astronomaly: Protege anomaly detection framework to identify the final candidates. Using a human-labelled set, we evaluate our pipeline on high-resolution (~7''), convolved (15''), and combined-feature MGCLS datasets. Interestingly, the high-resolution features identify diffuse sources more efficiently than the convolved resolution, which are in turn outperformed by the combined features. Of the top 100 sources ranked by Protege, 99% exhibit diffuse characteristics, with 55% confirmed as cluster-related emission. Our work shows that Protege can identify diffuse emission with minimal human labelling effort, offering a powerful, scalable tool capable of detecting both known and novel diffuse radio sources.

astro-ph.IM

A Guided Unconditional Diffusion Model to Synthesize and Inpaint Radio Galaxies from FIRST, MGCLS and Radio Zoo

We present a masked-guided approach for a denoising diffusion probabilistic model (DDPM) trained to generate and inpaint realistic radio galaxy images. The inpainting capability is particularly relevant for reconstructing incomplete observations, improving downstream tasks such as source characterization and morphological classification. We train the model on a combination of the FIRST survey, Radio Galaxy Zoo, and cutouts from the MGCLS survey, enabling it to capture a broad range of radio galaxy morphologies across different observational regimes. We evaluate the realism of the generated samples through statistical comparisons with real data, ensuring consistency in key morphological and intensity distributions. Our unconditional model produces morphologically plausible galaxies while maintaining diversity, highlighting the suitability of diffusion models for this task. This approach provides a scalable alternative to computationally expensive simulations and enables effective data augmentation for machine learning applications in radio astronomy, including source detection, classification, and image reconstruction.

astro-ph.GA

TEGLIE: Transformer encoders as strong gravitational lens finders in KiDS

We apply a state-of-the-art transformer algorithm to 221 deg$^2$ of the Kilo Degree Survey (KiDS) to search for new strong gravitational lenses (SGL). We test four transformer encoders trained on simulated data from the Strong Lens Finding Challenge on KiDS survey data. The best performing model is fine-tuned on real images of SGL candidates identified in previous searches. To expand the dataset for fine-tuning, data augmentation techniques are employed, including rotation, flipping, transposition, and white noise injection. The network fine-tuned with rotated, flipped, and transposed images exhibited the best performance and is used to hunt for SGL in the overlapping region of the Galaxy And Mass Assembly (GAMA) and KiDS surveys on galaxies up to $z$=0.8. Candidate SGLs are matched with those from other surveys and examined using GAMA data to identify blended spectra resulting from the signal from multiple objects in a fiber. We observe that fine-tuning the transformer encoder to the KiDS data reduces the number of false positives by 70%. Additionally, applying the fine-tuned model to a sample of $\sim$ 5,000,000 galaxies results in a list of $\sim$ 51,000 SGL candidates. Upon visual inspection, this list is narrowed down to 231 candidates. Combined with the SGL candidates identified in the model testing, our final sample includes 264 candidates, with 71 high-confidence SGLs of which 44 are new discoveries. We propose fine-tuning via real augmented images as a viable approach to mitigating false positives when transitioning from simulated lenses to real surveys. Additionally, we provide a list of 121 false positives that exhibit features similar to lensed objects, which can benefit the training of future machine learning models in this field.

astro-ph.GA

Astronomaly at scale: searching for anomalies amongst 4 million galaxies

Modern astronomical surveys are producing datasets of unprecedented size and richness, increasing the potential for high-impact scientific discovery. This possibility, coupled with the challenge of exploring a large number of sources, has led to the development of novel machine-learning-based anomaly detection approaches, such as Astronomaly. For the first time, we test the scalability of Astronomaly by applying it to almost 4 million images of galaxies from the Dark Energy Camera Legacy Survey. We use a trained deep learning algorithm to learn useful representations of the images and pass these to the anomaly detection algorithm isolation forest, coupled with Astronomaly's active learning method, to discover interesting sources. We find that data selection criteria have a significant impact on the trade-off between finding rare sources such as strong lenses and introducing artefacts into the dataset. We demonstrate that active learning is required to identify the most interesting sources and reduce artefacts, while anomaly detection methods alone are insufficient. Using Astronomaly, we find 1635 anomalies among the top 2000 sources in the dataset after applying active learning, including eight strong gravitational lens candidates, 1609 galaxy merger candidates, and 18 previously unidentified sources exhibiting highly unusual morphology. Our results show that by leveraging the human-machine interface, Astronomaly is able to rapidly identify sources of scientific interest even in large datasets.

astro-ph.IM

Practical Galaxy Morphology Tools from Deep Supervised Representation Learning

Astronomers have typically set out to solve supervised machine learning problems by creating their own representations from scratch. We show that deep learning models trained to answer every Galaxy Zoo DECaLS question learn meaningful semantic representations of galaxies that are useful for new tasks on which the models were never trained. We exploit these representations to outperform several recent approaches at practical tasks crucial for investigating large galaxy samples. The first task is identifying galaxies of similar morphology to a query galaxy. Given a single galaxy assigned a free text tag by humans (e.g. "#diffuse"), we can find galaxies matching that tag for most tags. The second task is identifying the most interesting anomalies to a particular researcher. Our approach is 100% accurate at identifying the most interesting 100 anomalies (as judged by Galaxy Zoo 2 volunteers). The third task is adapting a model to solve a new task using only a small number of newly-labelled galaxies. Models fine-tuned from our representation are better able to identify ring galaxies than models fine-tuned from terrestrial images (ImageNet) or trained from scratch. We solve each task with very few new labels; either one (for the similarity search) or several hundred (for anomaly detection or fine-tuning). This challenges the longstanding view that deep supervised methods require new large labelled datasets for practical use in astronomy. To help the community benefit from our pretrained models, we release our fine-tuning code Zoobot. Zoobot is accessible to researchers with no prior experience in deep learning.

astro-ph.GA