SearcharxivSearch

arXiv subjects

Hugh Dickinson

Publications and source records attributed to Hugh Dickinson.

At least 19 recordsLinked to original sources

The fraction of clumpy star-forming galaxies in nearby galaxies from CLAUDS and HSC-SSP

Massive, star-forming clumps are regions of intensive star-formation that are commonly observed in high-redshift ($z > 1$) galaxies. Observations of low-redshift clumpy galaxy analogues are rare but the availability of wide-field galaxy survey data makes the detection of large clumpy galaxy samples much more feasible. We present a population of 12,790 star-forming clumps detected in a mass-complete sample of 5,395 star-forming galaxies (SFGs) at redshifts $z\leq0.32$, located in the XMM-LSS, E-COSMOS and DEEP2-3 fields observed by the Hyper Suprime-Cam Subaru Strategic Survey (HSC-SSP) and CFHT Large Area U-band Deep Survey (CLAUDS). The clumps were detected using an improved version of our Deep Learning (DL)-based object detection framework which uses the ZOOBOT foundation DL-model as a 'backbone' feature extractor. We determined the fraction of star-forming galaxies hosting at least one off-centre clump ($f_{\mathrm{clumpy}}$) based on a clump definition that requires a clump-galaxy flux ratio in the CLAUDS u-band of $\geq8\%$. We estimate $f_{\mathrm{clumpy}}$ to decrease from $\sim$31\% at $z\sim0.3$ to $\sim$23\% at $z \sim 0.1$, which aligns well with a low-redshift extrapolation of the clumpy fraction that is measured using high-redshift observations. At fixed redshift, $f_{\mathrm{clumpy}}$ is negatively correlated with the stellar mass and positively correlated with the specific star-formation rate (sSFR) of the host galaxies. When the clump definition is changed to include only clumps with a stellar mass of $M_{\mathrm{cl}} \geq 10^7 M_\odot$, we observe a highly increased clumpy fraction of $\sim$60\% that tends to increase with the stellar mass of the host galaxies but does not show a dependence on the sSFR of the host galaxies.

astro-ph.GA

Star-forming clump detection in nearby galaxies using Faster R-CNN and $ugrizy$ imaging data from CLAUDS and HSC-SSP

Giant Star-forming Clumps (GSFCs) are kpc-scale regions of enhanced star-formation with stellar masses of $10^7$ to $10^9\,M_\odot$ that are commonly observed in high-redshift galaxies but are rarely detected in low-redshift ($z\lesssim0.5$) galaxy analogues. However, the availability of wide-field galaxy survey data makes it possible to identify potential star-forming clumps in large samples of low-redshift galaxies using object detection models that are based on Deep Learning (DL) techniques. We apply a novel DL-based object detection model to galaxies observed by the Hyper Suprime-Cam Subaru Strategic Survey (HSC-SSP) and CFHT Large Area U-band Deep Survey (CLAUDS). Our model is based on the the Faster Region-Based Convolutional Neural Network (Faster R-CNN or FRCNN) object detection framework but expanded to process the six $ugrizy$ filter band images simultaneously and identify not only clumps and their locations in the host galaxy but also additional contaminants. By adopting the \textsc{Zoobot} foundation DL-model as a feature extraction backbone, we also demonstrate one of the first applications of \textsc{Zoobot} in a downstream task for object detection. Our model achieves a detection completeness of $\gtrsim 0.9$ and purity of $\gtrsim 0.8$ which were validated on a large set of real galaxies into which simulated clumps were injected.

astro-ph.IM

Citizen Science Research with the Square Kilometre Array Observatory (SKAO)

Over the past two decades, internet-enabled citizen science research (CSR) has contributed to significant discoveries while involving millions of people in the research process. Our review highlights CSR in extragalactic radio astronomy and emphasises that such approaches will become increasingly relevant across radio astronomy in the era of the Square Kilometre Array (SKA). As astronomical data volumes grow, CSR is converging with Artificial Intelligence and Machine Learning (AI/ML), creating hybrid human-machine frameworks suited to big-data challenges. Two CSR platforms, Radio Galaxy Zoo and RAD@home, demonstrate success: the former excels in large-scale, web-based catalogue creation, while the latter combines structured training with collaborative discovery. Following this, we propose CSR with the SKA, namely SKA@home, with two modes: one purely web-based and the other in collaboratory mode with national training programmes. We argue that CSR can complement, and at times surpass, automated AI/ML pipelines, particularly in identifying rare, intricate, or unexpected features. Illustrative CSR discoveries include an episodic wide-angle-tailed radio galaxy, a jet-galaxy interaction, a collimated synchrotron thread, a twin-ring odd radio circle, and a large-scale shock ahead of a cluster-infalling galaxy. Consistent with the IAU's recognition of CSR as a driver of Astronomy for Development and the United Nations' affirmation of participation in science as a universal human right, both the SKA construction proposal and outreach strategies show commitment to enabling CSR with SKA. The proposed SKA@home would not only enhance the early discovery potential of SKA data but also initiate a deeper and more meaningful connection with society at large.

astro-ph.GA

Spatio-Spectroscopic Representation Learning using Unsupervised Convolutional Long-Short Term Memory Networks

Integral Field Spectroscopy (IFS) surveys offer a unique new landscape in which to learn in both spatial and spectroscopic dimensions and could help uncover previously unknown insights into galaxy evolution. In this work, we demonstrate a new unsupervised deep learning framework using Convolutional Long-Short Term Memory Network Autoencoders to encode generalized feature representations across both spatial and spectroscopic dimensions spanning $19$ optical emission lines (3800A $< \lambda <$ 8000A) among a sample of $\sim 9000$ galaxies from the MaNGA IFS survey. As a demonstrative exercise, we assess our model on a sample of $290$ Active Galactic Nuclei (AGN) and highlight scientifically interesting characteristics of some highly anomalous AGN.

astro-ph.GA

Galaxy Zoo Evo: 1 million human-annotated images of galaxies

We introduce Galaxy Zoo Evo, a labeled dataset for building and evaluating foundation models on images of galaxies. GZ Evo includes 104M crowdsourced labels for 823k images from four telescopes. Each image is labeled with a series of fine-grained questions and answers (e.g. "featured galaxy, two spiral arms, tightly wound, merging with another galaxy"). These detailed labels are useful for pretraining or finetuning. We also include four smaller sets of labels (167k galaxies in total) for downstream tasks of specific interest to astronomers, including finding strong lenses and describing galaxies from the new space telescope Euclid. We hope GZ Evo will serve as a real-world benchmark for computer vision topics such as domain adaption (from terrestrial to astronomical, or between telescopes) or learning under uncertainty from crowdsourced labels. We also hope it will support a new generation of foundation models for astronomy; such models will be critical to future astronomers seeking to better understand our universe.

astro-ph.IM

Galaxy Zoo: Cosmic Dawn -- morphological classifications for over 41,000 galaxies in the Euclid Deep Field North from the Hawaii Two-0 Cosmic Dawn survey

We present morphological classifications of over 41,000 galaxies out to $z_{\rm phot}\sim2.5$ across six square degrees of the Euclid Deep Field North (EDFN) from the Hawaii Twenty Square Degree (H20) survey, a part of the wider Cosmic Dawn survey. Galaxy Zoo citizen scientists play a crucial role in the examination of large astronomical data sets through crowdsourced data mining of extragalactic imaging. This iteration, Galaxy Zoo: Cosmic Dawn (GZCD), saw tens of thousands of volunteers and the deep learning foundation model Zoobot collectively classify objects in ultra-deep multiband Hyper Suprime-Cam (HSC) imaging down to a depth of $m_{HSC-i} = 21.5$. Here, we present the details and general analysis of this iteration, including the use of Zoobot in an active learning cycle to improve both model performance and volunteer experience, as well as the discovery of 51 new gravitational lenses in the EDFN. We also announce the public data release of the classifications for over 45,000 subjects, including more than 41,000 galaxies (median $z_{\rm phot}$ of $0.42\pm0.23$), along with their associated image cutouts. This data set provides a valuable opportunity for follow-up imaging of objects in the EDFN as well as acting as a truth set for training deep learning models for application to ground-based surveys like that of the Ultraviolet Near-Infrared Optical Northern Survey (UNIONS) collaboration and the newly operational Vera C. Rubin Observatory.

astro-ph.GA

Overview of the ESCAPE Dark Matter Test Science Project for Astronomers

The search for dark matter has been ongoing for decades within both astrophysics and particle physics. Both fields have employed different approaches and conceived a variety of methods for constraining the properties of dark matter, but have done so in relative isolation of one another. From an astronomer's perspective, it can be challenging to interpret the results of dark matter particle physics experiments and how these results apply to astrophysical scales. Over the past few years, the ESCAPE Dark Matter Test Science Project has been developing tools to aid the particle physics community in constraining dark matter properties; however, ESCAPE itself also aims to foster collaborations between research disciplines. This is especially important in the search for dark matter, as while particle physics is concerned with detecting the particles themselves, all of the evidence for its existence lies solely within astrophysics and cosmology. Here, we present a short review of the progress made by the Dark Matter Test Science Project and their applications to existing experiments, with a view towards how this project can foster complementary with astrophysical observations.

astro-ph.CO

Designing cultured tissue moulds using evolutionary strategies

There is an unmet need for artificial intelligence techniques that can speed up the design of growth strategies for cultured tissues. Cultured tissue is increasingly important for a range of applications such as cultivated meat, pharmaceutical assays and regenerative medicine. In this paper, we introduce a method based around evolutionary strategies, machine learning and biophysical simulations that can be used to speed up the process of identifying new tissue growth strategies for these diverse applications. We demonstrate the method by designing tethering strategies to grow tissues containing various cell types with desirable properties such as high cellular alignment and uniform density.

physics.bio-ph

Galaxy Zoo CEERS: Bar fractions up to z~4.0

We study the evolution of the bar fraction in disc galaxies between $0.5 < z < 4.0$ using multi-band coloured images from JWST CEERS. These images were classified by citizen scientists in a new phase of the Galaxy Zoo project called GZ CEERS. Citizen scientists were asked whether a strong or weak bar was visible in the host galaxy. After considering multiple corrections for observational biases, we find that the bar fraction decreases with redshift in our volume-limited sample (n = 398); from $25^{+6}_{-4}$% at $0.5 < z < 1.0$ to $3^{+6}_{-1}$% at $3.0 < z < 4.0$. However, we argue it is appropriate to interpret these fractions as lower limits. Disentangling real changes in the bar fraction from detection biases remains challenging. Nevertheless, we find a significant number of bars up to $z = 2.5$. This implies that discs are dynamically cool or baryon-dominated, enabling them to host bars. This also suggests that bar-driven secular evolution likely plays an important role at higher redshifts. When we distinguish between strong and weak bars, we find that the weak bar fraction decreases with increasing redshift. In contrast, the strong bar fraction is constant between $0.5 < z < 2.5$. This implies that the strong bars found in this work are robust long-lived structures, unless the rate of bar destruction is similar to the rate of bar formation. Finally, our results are consistent with disc instabilities being the dominant mode of bar formation at lower redshifts, while bar formation through interactions and mergers is more common at higher redshifts.

astro-ph.GA

Rapid prediction of organisation in engineered corneal, glial and fibroblast tissues using machine learning and biophysical models

We present a machine learning approach for predicting the organisation of corneal, glial and fibroblast cells in 3D cultures used for tissue engineering. Our machine-learning-based method uses a powerful generative adversarial network architecture called pix2pix, which we train using results from biophysical contractile network dipole orientation (CONDOR) simulations. In the following, we refer to the machine learning method as the RAPTOR (RApid Prediction of Tissue ORganisation) approach. A training data set containing a range of CONDOR simulations is created, covering a range of underlying model parameters. Validation of the trained neural network is carried out by comparing predictions with cultured glial, corneal, and fibroblast tissues, with good agreements for both CONDOR and RAPTOR approaches. An approach is developed to determine CONDOR model parameters for specific tissues using a fit to tissue properties. RAPTOR outputs a variety of tissue properties, including cell densities of cell alignments and tension. Since it is fast, it could be valuable for the design of tethered moulds for tissue growth.

physics.bio-ph

Citizen Science in European Research Infrastructures

Major European Union-funded research infrastructure and open science projects have traditionally included dissemination work, for mostly one-way communication of the research activities. Here we present and review our radical re-envisioning of this work, by directly engaging citizen science volunteers into the research. We summarise the citizen science in the Horizon-funded projects ASTERICS (Astronomy ESFRI and Research Infrastructure Clusters) and ESCAPE (European Science Cluster of Astronomy and Particle Physics ESFRI Research Infrastructures), engaging hundreds of thousands of volunteers in providing millions of data mining classifications. Not only does this have enormously more scientific and societal impact than conventional dissemination, but it facilitates the direct research involvement of what is often arguably the most neglected stakeholder group in Horizon projects, the science-inclined public. We conclude with recommendations and opportunities for deploying crowdsourced data mining in the physical sciences, noting that the primary goal is always the fundamental research question; if public engagement is the primary goal to optimise, then other, more targeted approaches may be more effective.

astro-ph.IM

Scaling Laws for Galaxy Images

We present the first systematic investigation of supervised scaling laws outside of an ImageNet-like context - on images of galaxies. We use 840k galaxy images and over 100M annotations by Galaxy Zoo volunteers, comparable in scale to Imagenet-1K. We find that adding annotated galaxy images provides a power law improvement in performance across all architectures and all tasks, while adding trainable parameters is effective only for some (typically more subjectively challenging) tasks. We then compare the downstream performance of finetuned models pretrained on either ImageNet-12k alone vs. additionally pretrained on our galaxy images. We achieve an average relative error rate reduction of 31% across 5 downstream tasks of scientific interest. Our finetuned models are more label-efficient and, unlike their ImageNet-12k-pretrained equivalents, often achieve linear transfer performance equal to that of end-to-end finetuning. We find relatively modest additional downstream benefits from scaling model size, implying that scaling alone is not sufficient to address our domain gap, and suggest that practitioners with qualitatively different images might benefit more from in-domain adaption followed by targeted downstream labelling.

cs.CV

Transfer learning for galaxy feature detection: Finding Giant Star-forming Clumps in low redshift galaxies using Faster R-CNN

Giant Star-forming Clumps (GSFCs) are areas of intensive star-formation that are commonly observed in high-redshift (z>1) galaxies but their formation and role in galaxy evolution remain unclear. High-resolution observations of low-redshift clumpy galaxy analogues are rare and restricted to a limited set of galaxies but the increasing availability of wide-field galaxy survey data makes the detection of large clumpy galaxy samples increasingly feasible. Deep Learning, and in particular CNNs, have been successfully applied to image classification tasks in astrophysical data analysis. However, one application of DL that remains relatively unexplored is that of automatically identifying and localising specific objects or features in astrophysical imaging data. In this paper we demonstrate the feasibility of using Deep learning-based object detection models to localise GSFCs in astrophysical imaging data. We apply the Faster R-CNN object detection framework (FRCNN) to identify GSFCs in low redshift (z<0.3) galaxies. Unlike other studies, we train different FRCNN models not on simulated images with known labels but on real observational data that was collected by the Sloan Digital Sky Survey Legacy Survey and labelled by volunteers from the citizen science project `Galaxy Zoo: Clump Scout'. The FRCNN model relies on a CNN component as a `backbone' feature extractor. We show that CNNs, that have been pre-trained for image classification using astrophysical images, outperform those that have been pre-trained on terrestrial images. In particular, we compare a domain-specific CNN -`Zoobot' - with a generic classification backbone and find that Zoobot achieves higher detection performance and also requires smaller training data sets to do so. Our final model is capable of producing GSFC detections with a completeness and purity of >=0.8 while only being trained on ~5,000 galaxy images.

astro-ph.GA

Using cGANs for Anomaly Detection: Identifying Astronomical Anomalies in JWST NIRcam Imaging

We present a proof of concept for mining JWST imaging data for anomalous galaxy populations using a conditional Generative Adversarial Network (cGAN). We train our model to predict long wavelength NIRcam fluxes (LW: F277W, F356W, F444W between 2.4 to 5.0μm) from short wavelength fluxes (SW: F115W, F150W, F200W between 0.6 to 2.3μm) in approximately 2000 galaxies. We test the cGAN on a population of 37 Extremely Red Objects (EROs) discovered by the CEERS JWST Team arXiv:2305.14418. Despite their red long wavelength colours, the EROs have blue short wavelength colours (F150W \- F200W equivalently 0 mag) indicative of bimodal SEDs. Surprisingly, given their unusual SEDs, we find that the cGAN accurately predicts the LW NIRcam fluxes of the EROs. However, it fails to predict LW fluxes for other rare astronomical objects, such as a merger between two galaxies, suggesting that the cGAN can be used to detect some anomalies

astro-ph.CO

Galaxy Zoo DESI: Detailed Morphology Measurements for 8.7M Galaxies in the DESI Legacy Imaging Surveys

We present detailed morphology measurements for 8.67 million galaxies in the DESI Legacy Imaging Surveys (DECaLS, MzLS, and BASS, plus DES). These are automated measurements made by deep learning models trained on Galaxy Zoo volunteer votes. Our models typically predict the fraction of volunteers selecting each answer to within 5-10\% for every answer to every GZ question. The models are trained on newly-collected votes for DESI-LS DR8 images as well as historical votes from GZ DECaLS. We also release the newly-collected votes. Extending our morphology measurements outside of the previously-released DECaLS/SDSS intersection increases our sky coverage by a factor of 4 (5,000 to 19,000 deg$^2$) and allows for full overlap with complementary surveys including ALFALFA and MaNGA.

astro-ph.GA

High-throughput design of cultured tissue moulds using a biophysical model

The technique presented here identifies tethered mould designs, optimised for growing cultured tissue with very highly-aligned cells. It is based on a microscopic biophysical model for polarised cellular hydrogels. There is an unmet need for tools to assist mould and scaffold designs for the growth of cultured tissues with bespoke cell organisations, that can be used in applications such as regenerative medicine, drug screening and cultured meat. High-throughput biophysical calculations were made for a wide variety of computer-generated moulds, with cell-matrix interactions and tissue-scale forces simulated using a contractile-network dipole-orientation model. Elongated moulds with central broadening and one of the following tethering strategies are found to lead to highly-aligned cells: (1) tethers placed within the bilateral protrusions resulting from an indentation on the short edge, to guide alignment (2) tethers placed within a single vertex to shrink the available space for misalignment. As such, proof-of-concept has been shown for mould and tethered scaffold design based on a recently developed biophysical model. The approach is applicable to a broad range of cell types that align in tissues and is extensible for 3D scaffolds.

physics.bio-ph

Rapid prediction of lab-grown tissue properties using deep learning

The interactions between cells and the extracellular matrix are vital for the self-organisation of tissues. In this paper we present proof-of-concept to use machine learning tools to predict the role of this mechanobiology in the self-organisation of cell-laden hydrogels grown in tethered moulds. We develop a process for the automated generation of mould designs with and without key symmetries. We create a large training set with $N=6500$ cases by running detailed biophysical simulations of cell-matrix interactions using the contractile network dipole orientation (CONDOR) model for the self-organisation of cellular hydrogels within these moulds. These are used to train an implementation of the \texttt{pix2pix} deep learning model, reserving $740$ cases that were unseen in the training of the neural network for training and validation. Comparison between the predictions of the machine learning technique and the reserved predictions from the biophysical algorithm show that the machine learning algorithm makes excellent predictions. The machine learning algorithm is significantly faster than the biophysical method, opening the possibility of very high throughput rational design of moulds for pharmaceutical testing, regenerative medicine and fundamental studies of biology. Future extensions for scaffolds and 3D bioprinting will open additional applications.

q-bio.TO

Galaxy Zoo: Clump Scout -- Design and first application of a two-dimensional aggregation tool for citizen science

Galaxy Zoo: Clump Scout is a web-based citizen science project designed to identify and spatially locate giant star forming clumps in galaxies that were imaged by the Sloan Digital Sky Survey Legacy Survey. We present a statistically driven software framework that is designed to aggregate two-dimensional annotations of clump locations provided by multiple independent Galaxy Zoo: Clump Scout volunteers and generate a consensus label that identifies the locations of probable clumps within each galaxy. The statistical model our framework is based on allows us to assign false-positive probabilities to each of the clumps we identify, to estimate the skill levels of each of the volunteers who contribute to Galaxy Zoo: Clump Scout and also to quantitatively assess the reliability of the consensus labels that are derived for each subject. We apply our framework to a dataset containing 3,561,454 two-dimensional points, which constitute 1,739,259 annotations of 85,286 distinct subjects provided by 20,999 volunteers. Using this dataset, we identify 128,100 potential clumps distributed among 44,126 galaxies. This dataset can be used to study the prevalence and demographics of giant star forming clumps in low-redshift galaxies. The code for our aggregation software framework is publicly available at: https://github.com/ou-astrophysics/BoxAggregator

astro-ph.GA