SearcharxivSearch

arXiv subjects

William J. Pearson

Publications and source records attributed to William J. Pearson.

12 recordsLinked to original sources

Performance of morphological classifiers for galaxy mergers compared to current machine learning methods

Aims. Non-parametric morphological statistics can be used for efficient classification of galaxy mergers. This work aims to compare the performance of morphological merger classifiers to state-of-the-art machine learning (ML) models. A secondary aim is to produce updated criteria for mergers based on non-parametric morphological statistics. Methods. The Gini coefficient (G), $M_{20}$ statistic, and concentration ($C$) were calculated for mock Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) images based on the IllustrisTNG and Horizon-AGN simulations, and observations from HSC-SSP. The IllustrisTNG images were used to find the line which best separates mergers and non-mergers in 2D morphological space with a Markov Chain Monte-Carlo (MCMC) method. Results. Based on the MCMC results, we classified galaxies with $G>(-0.267\pm0.081)M_{20}+(0.143\pm0.012)$ or $G>(0.162\pm0.048)C-(0.149\pm0.12)$ as mergers, these criteria had precisions of 69.5\% and 72.3\% respectively when applied to previously unseen IllustrisTNG mock HSC-SSP images. The precisions of the morphological classifications are consistent with state-of-the-art ML methods. The morphological classifiers were found to be effective at selecting only pre-mergers; post-merger galaxies are indistinguishable from non-mergers in terms of their $G$, $M_{20}$, and $C$ values. Morphological classifiers displayed a similar robustness to new data to ML methods up to a redshift of $\sim0.52$ and maintained robustness better than ML methods based on convolutional neural networks in the redshift range $0.52<z<1$. Conclusions. This work presents updated morphological classifiers which achieve similar precisions to ML based merger classifiers with a high robustness to new data. New morphological statistics are needed to identify the features of post-merger galaxies.

astro-ph.GA

From DES to KiDS: Domain adaptation for cross-survey detection of low-surface-brightness galaxies

Low-surface-brightness galaxies (LSBGs) are vital for understanding galaxy formation, but their diffuse nature makes them challenging to detect. Upcoming large-scale surveys are expected to uncover large numbers of LSBGs, requiring robust automated methods to identify them across heterogeneous datasets. As a precursor to the Legacy Survey of Space and Time (LSST) and Euclid, we explore domain adaptation techniques for cross-survey LSBG identification. Using models trained on the Dark Energy Survey (DES), we search for LSBGs in the Kilo-Degree Survey Data Release 5 (KiDS DR5). We used an ensemble consisting of one convolutional neural network (CNN) and two transformer models trained on DES cutouts and applied to KiDS DR5 imaging data. Structural parameters were estimated with galfitm, and photometric redshifts and stellar population properties were estimated through spectral energy distribution fitting with CIGALE. We identify 20,180 LSBGs and 434 ultra-diffuse galaxies (UDGs) in KiDS DR5. Their structural parameters are similar to known LSBGs from DES and the Hyper Suprime-Cam SSP Survey (HSC-SSP). The KiDS-LSBGs follow a continuous size-luminosity relation connecting classical dwarf galaxies and UDGs, and their colours are bimodal ($\sim73\%$ blue, $\sim27\%$ red). Cross-matching with spectroscopic and cluster catalogues provides redshifts for 4,913 systems, enabling a systematic characterisation of the star-forming main sequence of LSBGs. Strong environmental trends are evident, with cluster LSBGs and UDGs exhibiting redder colours and reduced star formation compared to non-cluster systems. We demonstrate that domain adaptation enables robust cross-survey LSBG identification with deep learning models, providing a scalable pathway for constructing homogeneous LSBG catalogues for the LSST and Euclid era.

astro-ph.GA

statmorph-lsst: Quantifying and correcting morphological biases in galaxy surveys

Quantitative morphology provides a key probe of galaxy evolution across cosmic time and environments. However, these metrics can be biased by changes in imaging quality - resolution and depth - either across the survey area or the sample. To prepare for the upcoming Rubin LSST data, we investigate this bias for all metrics measured by statmorph and single-component Sérsic fitting with Galfit. We find that geometrical measurements (ellipticity, axis ratio, Petrosian radius, and effective radius) are robust within 10% at most depths and resolutions. Light concentration measurements ($C$, Gini, $M_{20}$) systematically decrease with resolution, leading low-mass or high-redshift bulge-dominated sources to appear indistinguishable from disks. Sérsic index $n$, while unbiased, suffers from a 20-40% uncertainty due to degeneracies in the Sérsic fit. Disturbance measurements ($A$, $A_S$, $D$) depend on signal-to-noise and are thus affected by noise and surface-brightness dimming. We quantify this dependence for each parameter, offer empirical correction functions, and show that the evolution in $C$ observed in JWST galaxies can be explained purely by observational biases. We propose two new measurements - isophotal asymmetry $A_X$ and substructure $St$ - that aim to resolve some of these biases. Finally, we provide a Python package statmorph-lsst implementing these changes and a full dataset that enables tests of custom functions (see text for links).

astro-ph.GA

Galaxy Zoo: Cosmic Dawn -- morphological classifications for over 41,000 galaxies in the Euclid Deep Field North from the Hawaii Two-0 Cosmic Dawn survey

We present morphological classifications of over 41,000 galaxies out to $z_{\rm phot}\sim2.5$ across six square degrees of the Euclid Deep Field North (EDFN) from the Hawaii Twenty Square Degree (H20) survey, a part of the wider Cosmic Dawn survey. Galaxy Zoo citizen scientists play a crucial role in the examination of large astronomical data sets through crowdsourced data mining of extragalactic imaging. This iteration, Galaxy Zoo: Cosmic Dawn (GZCD), saw tens of thousands of volunteers and the deep learning foundation model Zoobot collectively classify objects in ultra-deep multiband Hyper Suprime-Cam (HSC) imaging down to a depth of $m_{HSC-i} = 21.5$. Here, we present the details and general analysis of this iteration, including the use of Zoobot in an active learning cycle to improve both model performance and volunteer experience, as well as the discovery of 51 new gravitational lenses in the EDFN. We also announce the public data release of the classifications for over 45,000 subjects, including more than 41,000 galaxies (median $z_{\rm phot}$ of $0.42\pm0.23$), along with their associated image cutouts. This data set provides a valuable opportunity for follow-up imaging of objects in the EDFN as well as acting as a truth set for training deep learning models for application to ground-based surveys like that of the Ultraviolet Near-Infrared Optical Northern Survey (UNIONS) collaboration and the newly operational Vera C. Rubin Observatory.

astro-ph.GA

Machine learning classification of baseband data of CHIME FRBs

Fast Radio Bursts (FRBs) are bright millisecond radio pulses. Their origin is still unknown in the field of astronomy. A notable distinction among FRBs is that some sources repeat, while others appear to be non-repeating events. Interestingly, repeating FRBs tend to exhibit broader temporal widths and narrower spectral bandwidths compared to non-repeat events, suggesting they may arise from different physical mechanisms. However, current radio telescopes have limited coverage and sensitivity, which hinders a complete survey with continuous long-term monitoring. This issue makes it difficult to confirm repeat activity and potentially leads to misclassification of repeaters as non-repeaters; these are referred to as repeater candidates. To address this, machine learning techniques have emerged as a useful tool for classifying distinct FRB types in previous studies. In this study, we utilize the CHIME/FRB baseband catalog with three orders of magnitude better time resolution than the intensity catalog. Measured fluences are available in the baseband catalog, while only upper limits are reported in the intensity catalog. We apply machine learning to the baseband catalog to evaluate classification outcomes. We identify 15 repeater candidates among 122 non-repeating FRBs in the baseband catalog. Additionally, our classification identifies 31 sources previously categorized as repeater candidates as non-repeaters, highlighting a significant difference from the prior work. Of these repeater candidates, 14 overlap with previous findings, while 1 is newly identified in this work. Notably, one of our candidates was confirmed as a repeater by CHIME/FRB. Follow-up observations for the 14 candidates are highly encouraged.

astro-ph.HE

Automated quasar continuum estimation using neural networks: a comparative study of deep-learning architectures

Context. Ongoing and upcoming large spectroscopic surveys are drastically increasing the number of observed quasar spectra, requiring the development of fast and accurate automated methods to estimate spectral continua. Aims. This study evaluates the performance of three neural networks (NN) - an autoencoder, a convolutional NN (CNN), and a U-Net - in predicting quasar continua within the rest-frame wavelength range of $1020~\textÅ$ to $2000~\textÅ$. The ability to generalize and predict galaxy continua within the range of $3500~\textÅ$ to $5500~\textÅ$ is also tested. Methods. The performance of these architectures is evaluated using the absolute fractional flux error (AFFE) on a library of mock quasar spectra for the WEAVE survey, and on real data from the Early Data Release observations of the Dark Energy Spectroscopic Instrument (DESI) and the VIMOS Public Extragalactic Redshift Survey (VIPERS). Results. The autoencoder outperforms the U-Net, achieving a median AFFE of 0.009 for quasars. The best model also effectively recovers the Ly$α$ optical depth evolution in DESI quasar spectra. With minimal optimization, the same architectures can be generalized to the galaxy case, with the autoencoder reaching a median AFFE of 0.014 and reproducing the D4000n break in DESI and VIPERS galaxies.

astro-ph.GA

Classifying merger stages with adaptive deep learning and cosmological hydrodynamical simulations

Hierarchical merging of galaxies plays an important role in galaxy formation and evolution. Mergers could trigger key evolutionary phases such as starburst activities and active accretion periods onto supermassive black holes at the centres of galaxies. We aim to detect mergers and merger stages (pre- and post-mergers) across cosmic history and test whether it is better to detect mergers and their merger stages simultaneously or hierarchically. In addition, we want to test the impact of merger time relative to the coalescence of merging galaxies. First, we generated realistic mock JWST images of simulated galaxies selected from the IllustrisTNG cosmological hydrodynamical simulations. Then we trained deep learning (DL) models in the Zoobot Python package to classify galaxies into merging/non-merging galaxies and their merger stages. We used two different set-ups: (i) two-stage, in which we classify galaxies into mergers and non-mergers and then classify the mergers into pre-mergers and post-mergers, and (ii) one-stage, in which merger/non-merger and merger stages are classified simultaneously. We found that the one-stage classification set-up moderately outperforms the two-stage set-up, offering better overall accuracy and precision, particularly for the non-merger class. Pre-mergers can be classified with the highest precision in both set-ups, possibly due to the more recognisable merging features and the presence of merging companions. The image signal-to-noise ratio affects the performance of the DL classifiers, but not much after a certain threshold is crossed. Both precision and recall of the classifiers depend strongly on merger time, finding it more difficult to identify true mergers observed at stages that are more distant to coalescence. For pre-mergers, we recommend selecting mergers which will merge in the next 0.4 Gyrs, to achieve a good balance between precision and recall.

astro-ph.GA

Probabilistic and progressive deblended far-infrared and sub-millimetre point source catalogues I. Methodology and first application in the COSMOS field

Single-dish far-infrared (far-IR) and sub-millimetre (sub-mm) point source catalogues and their connections with catalogues at other wavelengths are of paramount importance. However, due to the large mismatch in spatial resolution, cross-matching galaxies at different wavelengths is challenging. This work aims to develop the next-generation deblended far-IR and sub-mm catalogues and present the first application in the COSMOS field. Our progressive deblending used the Bayesian probabilistic framework known as XID+. The deblending started from the Spitzer/MIPS 24 micron data, using an initial prior list composed of sources selected from the COSMOS2020 catalogue and radio catalogues from the VLA and the MeerKAT surveys, based on spectral energy distribution modelling which predicts fluxes of the known sources at the deblending wavelength. To speed up flux prediction, we made use of a neural network-based emulator. After deblending the 24 micron data, we proceeded to the Herschel PACS (100 & 160 micron) and SPIRE wavebands (250, 350 & 500 micron). Each time we constructed a tailor-made prior list based on the predicted fluxes of the known sources. Using simulated far-IR and sub-mm sky, we detailed the performance of our deblending pipeline. After validation with simulations, we then deblended the real observations from 24 to 500 micron and compared with blindly extracted catalogues and previous versions of deblended catalogues. As an additional test, we deblended the SCUBA-2 850 micron map and compared our deblended fluxes with ALMA measurements, which demonstrates a higher level of flux accuracy compared to previous results.We publicly release our XID+ deblended point source catalogues. These deblended long-wavelength data are crucial for studies such as deriving the fraction of dust-obscured star formation and better separation of quiescent galaxies from dusty star-forming galaxies.

astro-ph.GA

Hidden depths in the local Universe: The Stellar Stream Legacy Survey

Mergers and tidal interactions between massive galaxies and their dwarf satellites are a fundamental prediction of the Lambda-Cold Dark Matter cosmology. These events are thought to provide important observational diagnostics of nonlinear structure formation. Stellar streams in the Milky Way and Andromeda are spectacular evidence for ongoing satellite disruption. However, constructing a statistically meaningful sample of tidal streams beyond the Local Group has proven a daunting observational challenge, and the full potential for deepening our understanding of galaxy assembly using stellar streams has yet to be realised. Here we introduce the Stellar Stream Legacy Survey, a systematic imaging survey of tidal features associated with dwarf galaxy accretion around a sample of ~3100 nearby galaxies within z~0.02, including about 940 Milky Way analogues. Our survey exploits public deep imaging data from the DESI Legacy Imaging Surveys, which reach surface brightness as faint as ~29 mag/arcsec^2 in the r band. As a proof of concept of our survey, we report the detection and broad-band photometry of 24 new stellar streams in the local Universe. We discuss how these observations can yield new constraints on galaxy formation theory through comparison to mock observations from cosmological galaxy simulations. These tests will probe the present-day mass assembly rate of galaxies, the stellar populations and orbits of satellites, the growth of stellar halos and the resilience of stellar disks to satellite bombardment.

astro-ph.GA

Metallicity-PAH Relation of MIR-selected Star-forming Galaxies in AKARI North Ecliptic Pole-wide Survey

We investigate the variation in the mid-infrared spectral energy distributions of 373 low-redshift ($z<0.4$) star-forming galaxies, which reflects a variety of polycyclic aromatic hydrocarbon (PAH) emission features. The relative strength of PAH emission is parameterized as $q_\mathrm{PAH}$, which is defined as the mass fraction of PAH particles in the total dust mass. With the aid of continuous mid-infrared photometric data points covering 7-24$μ$m and far-infrared flux densities, $q_\mathrm{PAH}$ values are derived through spectral energy distribution fitting. The correlation between $q_\mathrm{PAH}$ and other physical properties of galaxies, i.e., gas-phase metallicity ($12+\mathrm{log(O/H)}$), stellar mass, and specific star-formation rate (sSFR) are explored. As in previous studies, $q_\mathrm{PAH}$ values of galaxies with high metallicity are found to be higher than those with low metallicity. The strength of PAH emission is also positively correlated with the stellar mass and negatively correlated with the sSFR. The correlation between $q_\mathrm{PAH}$ and each parameter still exists even after the other two parameters are fixed. In addition to the PAH strength, the application of metallicity-dependent gas-to-dust mass ratio appears to work well to estimate gas mass that matches the observed relationship between molecular gas and physical parameters. The result obtained will be used to calibrate the observed PAH luminosity-total infrared luminosity relation, based on the variation of MIR-FIR SED, which is used in the estimation of hidden star formation.

astro-ph.GA

Active galactic nuclei catalog from the AKARI NEP Wide field

Context. The North Ecliptic Pole (NEP) field provides a unique set of panchromatic data, well suited for active galactic nuclei (AGN) studies. Selection of AGN candidates is often based on mid-infrared (MIR) measurements. Such method, despite its effectiveness, strongly reduces a catalog volume due to the MIR detection condition. Modern machine learning techniques can solve this problem by finding similar selection criteria using only optical and near-infrared (NIR) data. Aims. Aims of this work were to create a reliable AGN candidates catalog from the NEP field using a combination of optical SUBARU/HSC and NIR AKARI/IRC data and, consequently, to develop an efficient alternative for the MIR-based AKARI/IRC selection technique. Methods. A set of supervised machine learning algorithms was tested in order to perform an efficient AGN selection. Best of the models were formed into a majority voting scheme, which used the most popular classification result to produce the final AGN catalog. Additional analysis of catalog properties was performed in form of the spectral energy distribution (SED) fitting via the CIGALE software. Results. The obtained catalog of 465 AGN candidates (out of 33 119 objects) is characterized by 73% purity and 64% completeness. This new classification shows consistency with the MIR-based selection. Moreover, 76% of the obtained catalog can be found only with the new method due to the lack of MIR detection for most of the new AGN candidates. Training data, codes and final catalog are available via the github repository. Final AGN candidates catalog will be also available via the CDS service after publication.

astro-ph.GA

Deep Learning for Galaxy Mergers in the Galaxy Main Sequence

Starburst galaxies are often found to be the result of galaxy mergers. As a result, galaxy mergers are often believed to lie above the galaxy main sequence: the tight correlation between stellar mass and star formation rate. Here, we aim to test this claim. Deep learning techniques are applied to images from the Sloan Digital Sky Survey to provide visual-like classifications for over 340 000 objects between redshifts of 0.005 and 0.1. The aim of this classification is to split the galaxy population into merger and non-merger systems and we are currently achieving an accuracy of 91.5%. Stellar masses and star formation rates are also estimated using panchromatic data for the entire galaxy population. With these preliminary data, the mergers are placed onto the full galaxy main sequence, where we find that merging systems lie across the entire star formation rate - stellar mass plane.

astro-ph.GA