SearcharxivSearch

arXiv subjects

Hongming Tang

Publications and source records attributed to Hongming Tang.

15 recordsLinked to original sources

Radio Galaxy Zoo EMU: Harnessing Citizen Science and AI to Advance Open Science Catalogues

Over the past decades, significant efforts have been devoted to developing sophisticated algorithms for automatically identifying and classifying radio sources in large surveys. However, even the most advanced methods face challenges in recognising complex radio structures and accurately associating radio emission with their host galaxies. Leveraging data from the ASKAP telescope and the Evolutionary Map of the Universe (EMU) survey, Radio Galaxy Zoo EMU (RGZ EMU) was created to generate high-quality radio source classifications for training deep learning models and cataloging millions of radio sources in the southern sky. By integrating novel machine learning techniques, including anomaly detection and natural language processing, our workflow actively engages citizen scientists to enhance classification accuracy. We present results from Phase I of the project and discuss how these data will contribute to improving open science catalogues like EMUCAT.

astro-ph.GA

Radio Galaxy Zoo: EMU -- paving the way for EMU cataloging using AI and citizen science

The Evolutionary Map of the Universe (EMU) survey with ASKAP is transforming our understanding of radio galaxies, AGN duty cycles, and cosmic structure. EMUCAT efficiently identifies compact radio sources, yet struggles with extended objects, requiring alternative approaches. The Radio Galaxy Zoo: EMU (RGZ EMU) project proposes a general framework that combines citizen science and machine learning to identify around 4 million extended sources in EMU. This framework is expected to enhance the EMUCAT cataloging on extended sources and can be further empowered with the introduction of cross-matched external data from surveys such as POSSUM and WALLABY.

astro-ph.IM

Category-based Galaxy Image Generation via Diffusion Models

Conventional galaxy generation methods rely on semi-analytical models and hydrodynamic simulations, which are highly dependent on physical assumptions and parameter tuning. In contrast, data-driven generative models do not have explicit physical parameters pre-determined, and instead learn them efficiently from observational data, making them alternative solutions to galaxy generation. Among these, diffusion models outperform Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) in quality and diversity. Leveraging physical prior knowledge to these models can further enhance their capabilities. In this work, we present GalCatDiff, the first framework in astronomy to leverage both galaxy image features and astrophysical properties in the network design of diffusion models. GalCatDiff incorporates an enhanced U-Net and a novel block entitled Astro-RAB (Residual Attention Block), which dynamically combines attention mechanisms with convolution operations to ensure global consistency and local feature fidelity. Moreover, GalCatDiff uses category embeddings for class-specific galaxy generation, avoiding the high computational costs of training separate models for each category. Our experimental results demonstrate that GalCatDiff significantly outperforms existing methods in terms of the consistency of sample color and size distributions, and the generated galaxies are both visually realistic and physically consistent. This framework will enhance the reliability of galaxy simulations and can potentially serve as a data augmentor to support future galaxy classification algorithm development.

astro-ph.IM

Galaxy stellar and total mass estimation using machine learning

Conventional galaxy mass estimation methods suffer from model assumptions and degeneracies. Machine learning, which reduces the reliance on such assumptions, can be used to determine how well present-day observations can yield predictions for the distributions of stellar and dark matter. In this work, we use a general sample of galaxies from the TNG100 simulation to investigate the ability of multi-branch convolutional neural network (CNN) based machine learning methods to predict the central (i.e., within $1-2$ effective radii) stellar and total masses, and the stellar mass-to-light ratio $M_*/L$. These models take galaxy images and spatially-resolved mean velocity and velocity dispersion maps as inputs. Such CNN-based models can in general break the degeneracy between baryonic and dark matter in the sense that the model can make reliable predictions on the individual contributions of each component. For example, with $r$-band images and two galaxy kinematic maps as inputs, our model predicting $M_*/L$ has a prediction uncertainty of 0.04 dex. Moreover, to investigate which (global) features significantly contribute to the correct predictions of the properties above, we utilize a gradient boosting machine. We find that galaxy luminosity dominates the prediction of all masses in the central regions, with stellar velocity dispersion coming next. We also investigate the main contributing features when predicting stellar and dark matter mass fractions ($f_*$, $f_{\rm DM}$) and the dark matter mass $M_{DM}$, and discuss the underlying astrophysics.

astro-ph.GA

Radio Galaxy Zoo: Leveraging latent space representations from variational autoencoder

We propose to learn latent space representations of radio galaxies, and train a very deep variational autoencoder (\protect\Verb+VDVAE+) on RGZ DR1, an unlabeled dataset, to this end. We show that the encoded features can be leveraged for downstream tasks such as classifying galaxies in labeled datasets, and similarity search. Results show that the model is able to reconstruct its given inputs, capturing the salient features of the latter. We use the latent codes of galaxy images, from MiraBest Confident and FR-DEEP NVSS datasets, to train various non-neural network classifiers. It is found that the latter can differentiate FRI from FRII galaxies achieving \textit{accuracy} $\ge 76\%$, \textit{roc-auc} $\ge 0.86$, \textit{specificity} $\ge 0.73$ and \textit{recall} $\ge 0.78$ on MiraBest Confident dataset, comparable to results obtained in previous studies. The performance of simple classifiers trained on FR-DEEP NVSS data representations is on par with that of a deep learning classifier (CNN based) trained on images in previous work, highlighting how powerful the compressed information is. We successfully exploit the learned representations to search for galaxies in a dataset that are semantically similar to a query image belonging to a different dataset. Although generating new galaxy images (e.g. for data augmentation) is not our primary objective, we find that the \protect\Verb+VDVAE+ model is a relatively good emulator. Finally, as a step toward detecting anomaly/novelty, a density estimator -- Masked Autoregressive Flow (\protect\Verb+MAF+) -- is trained on the latent codes, such that the log-likelihood of data can be estimated. The downstream tasks conducted in this work demonstrate the meaningfulness of the latent codes.

astro-ph.GA

A model local interpretation routine for deep learning based radio galaxy classification

Radio galaxy morphological classification is one of the critical steps when producing source catalogues for large-scale radio continuum surveys. While many recent studies attempted to classify source radio morphology from survey image data using deep learning algorithms (i.e., Convolutional Neural Networks), they concentrated on model robustness most time. It is unclear whether a model similarly makes predictions as radio astronomers did. In this work, we used Local Interpretable Model-agnostic Explanation (LIME), an state-of-the-art eXplainable Artificial Intelligence (XAI) technique to explain model prediction behaviour and thus examine the hypothesis in a proof-of-concept manner. In what follows, we describe how \textbf{LIME} generally works and early results about how it helped explain predictions of a radio galaxy classification model using this technique.

astro-ph.IM

Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy

We present a novel natural language processing (NLP) approach to deriving plain English descriptors for science cases otherwise restricted by obfuscating technical terminology. We address the limitations of common radio galaxy morphology classifications by applying this approach. We experimentally derive a set of semantic tags for the Radio Galaxy Zoo EMU (Evolutionary Map of the Universe) project and the wider astronomical community. We collect 8,486 plain English annotations of radio galaxy morphology, from which we derive a taxonomy of tags. The tags are plain English. The result is an extensible framework which is more flexible, more easily communicated, and more sensitive to rare feature combinations which are indescribable using the current framework of radio astronomy classifications.

astro-ph.GA

A New Task: Deriving Semantic Class Targets for the Physical Sciences

We define deriving semantic class targets as a novel multi-modal task. By doing so, we aim to improve classification schemes in the physical sciences which can be severely abstracted and obfuscating. We address this task for upcoming radio astronomy surveys and present the derived semantic radio galaxy morphology class targets.

astro-ph.IM

Identifying anomalous radio sources in the EMU Pilot Survey using a complexity-based approach

The Evolutionary Map of the Universe (EMU) large-area radio continuum survey will detect tens of millions of radio galaxies, giving an opportunity for the detection of previously unknown classes of objects. To maximise the scientific value and make new discoveries, the analysis of this data will need to go beyond simple visual inspection. We propose the coarse-grained complexity, a simple scalar quantity relating to the minimum description length of an image, that can be used to identify unusual structures. The complexity can be computed without reference to the broader sample or existing catalogue data, making the computation efficient on new surveys at very large scales (such as the full EMU survey). We apply our coarse-grained complexity measure to data from the EMU Pilot Survey to detect and confirm anomalous objects in this data set and produce an anomaly catalogue. Rather than work with existing catalogue data using a specific source detection algorithm, we perform a blind scan of the area, computing the complexity using a sliding square aperture. The effectiveness of the complexity measure for identifying anomalous objects is evaluated using crowd-sourced labels generated via the Zooniverse.org platform. We find that the complexity scan identifies unusual sources, such as odd radio circles, by partitioning on complexity. We achieve partitions where 5\% of the data is estimated to be 86\% complete, and 0.5\% is estimated to be 94\% pure, with respect to anomalies and use this to produce an anomaly catalogue.

astro-ph.IM

Radio Galaxy Zoo: Using semi-supervised learning to leverage large unlabelled data-sets for radio galaxy classification under data-set shift

In this work we examine the classification accuracy and robustness of a state-of-the-art semi-supervised learning (SSL) algorithm applied to the morphological classification of radio galaxies. We test if SSL with fewer labels can achieve test accuracies comparable to the supervised state-of-the-art and whether this holds when incorporating previously unseen data. We find that for the radio galaxy classification problem considered, SSL provides additional regularisation and outperforms the baseline test accuracy. However, in contrast to model performance metrics reported on computer science benchmarking data-sets, we find that improvement is limited to a narrow range of label volumes, with performance falling off rapidly at low label volumes. Additionally, we show that SSL does not improve model calibration, regardless of whether classification is improved. Moreover, we find that when different underlying catalogues drawn from the same radio survey are used to provide the labelled and unlabelled data-sets required for SSL, a significant drop in classification performance is observered, highlighting the difficulty of applying SSL techniques under dataset shift. We show that a class-imbalanced unlabelled data pool negatively affects performance through prior probability shift, which we suggest may explain this performance drop, and that using the Frechet Distance between labelled and unlabelled data-sets as a measure of data-set shift can provide a prediction of model performance, but that for typical radio galaxy data-sets with labelled sample volumes of O(1000), the sample variance associated with this technique is high and the technique is in general not sufficiently robust to replace a train-test cycle.

astro-ph.GA

Structured Variational Inference for Simulating Populations of Radio Galaxies

We present a model for generating postage stamp images of synthetic Fanaroff-Riley Class I and Class II radio galaxies suitable for use in simulations of future radio surveys such as those being developed for the Square Kilometre Array. This model uses a fully-connected neural network to implement structured variational inference through a variational auto-encoder and decoder architecture. In order to optimise the dimensionality of the latent space for the auto-encoder we introduce the radio morphology inception score (RAMIS), a quantitative method for assessing the quality of generated images, and discuss in detail how data pre-processing choices can affect the value of this measure. We examine the 2-dimensional latent space of the VAEs and discuss how this can be used to control the generation of synthetic populations, whilst also cautioning how it may lead to biases when used for data augmentation.

astro-ph.IM

Attention-gating for improved radio galaxy classification

In this work we introduce attention as a state of the art mechanism for classification of radio galaxies using convolutional neural networks. We present an attention-based model that performs on par with previous classifiers while using more than 50% fewer parameters than the next smallest classic CNN application in this field. We demonstrate quantitatively how the selection of normalisation and aggregation methods used in attention-gating can affect the output of individual models, and show that the resulting attention maps can be used to interpret the classification choices made by the model. We observe that the salient regions identified by the our model align well with the regions an expert human classifier would attend to make equivalent classifications. We show that while the selection of normalisation and aggregation may only minimally affect the performance of individual models, it can significantly affect the interpretability of the respective attention maps and by selecting a model which aligns well with how astronomers classify radio sources by eye, a user can employ the model in a more effective manner.

astro-ph.GA

Transfer learning for radio galaxy classification

In the context of radio galaxy classification, most state-of-the-art neural network algorithms have been focused on single survey data. The question of whether these trained algorithms have cross-survey identification ability or can be adapted to develop classification networks for future surveys is still unclear. One possible solution to address this issue is transfer learning, which re-uses elements of existing machine learning models for different applications. Here we present radio galaxy classification based on a 13-layer Deep Convolutional Neural Network (DCNN) using transfer learning methods between different radio surveys. We find that our machine learning models trained from a random initialization achieve accuracies comparable to those found elsewhere in the literature. When using transfer learning methods, we find that inheriting model weights pre-trained on FIRST images can boost model performance when re-training on lower resolution NVSS data, but that inheriting pre-trained model weights from NVSS and re-training on FIRST data impairs the performance of the classifier. We consider the implication of these results in the context of future radio surveys planned for next-generation radio telescopes such as ASKAP, MeerKAT, and SKA1-MID.

astro-ph.IM

Radio Galaxy Zoo: The Distortion of Radio Galaxies by Galaxy Clusters

We study the impact of cluster environment on the morphology of a sample of 4304 extended radio galaxies from Radio Galaxy Zoo. A total of 87% of the sample lies within a projected 15 Mpc of an optically identified cluster. Brightest cluster galaxies (BCGs) are more likely than other cluster members to be radio sources, and are also moderately bent. The surface density as a function of separation from cluster center of non-BCG radio galaxies follows a power law with index $-1.10\pm 0.03$ out to $10~r_{500}$ ($\sim 7~$Mpc), which is steeper than the corresponding distribution for optically selected galaxies. Non-BCG radio galaxies are statistically more bent the closer they are to the cluster center. Within the inner $1.5~r_{500}$ ($\sim 1~$Mpc) of a cluster, non-BCG radio galaxies are statistically more bent in high-mass clusters than in low-mass clusters. Together, we find that non-BCG sources are statistically more bent in environments that exert greater ram pressure. We use the orientation of bent radio galaxies as an indicator of galaxy orbits and find that they are preferentially in radial orbits. Away from clusters, there is a large population of bent radio galaxies, limiting their use as cluster locators; however, they are still located within statistically overdense regions. We investigate the asymmetry in the tail length of sources that have their tails aligned along the radius vector from the cluster center, and find that the length of the inward-pointing tail is weakly suppressed for sources close to the center of the cluster.

astro-ph.GA

Radio Galaxy Zoo: ClaRAN - A Deep Learning Classifier for Radio Morphologies

The upcoming next-generation large area radio continuum surveys can expect tens of millions of radio sources, rendering the traditional method for radio morphology classification through visual inspection unfeasible. We present ClaRAN - Classifying Radio sources Automatically with Neural networks - a proof-of-concept radio source morphology classifier based upon the Faster Region-based Convolutional Neutral Networks (Faster R-CNN) method. Specifically, we train and test ClaRAN on the FIRST and WISE images from the Radio Galaxy Zoo Data Release 1 catalogue. ClaRAN provides end users with automated identification of radio source morphology classifications from a simple input of a radio image and a counterpart infrared image of the same region. ClaRAN is the first open-source, end-to-end radio source morphology classifier that is capable of locating and associating discrete and extended components of radio sources in a fast (< 200 milliseconds per image) and accurate (>= 90 %) fashion. Future work will improve ClaRAN's relatively lower success rates in dealing with multi-source fields and will enable ClaRAN to identify sources on much larger fields without loss in classification accuracy.

astro-ph.IM