SearcharxivSearch

arXiv subjects

Alberto Krone-Martins

Publications and source records attributed to Alberto Krone-Martins.

At least 19 recordsLinked to original sources

Lightstack: A Python Package for Creating Photometric Data Cubes

Multi-band photometry traces diverse physical processes across a wide range of wavelengths. In recent decades, this field has been driven by the rapid growth of multi-imaging datasets, from high-resolution observation from Hubble Space Telescope and James Webb Space Telescope to the forthcoming large-scale surveys enabled by the Roman Space Telescope and Rubin Observatory, for example. In this work, we present lightstack, a Python package for combining standalone images into photometric data cubes. The workflow consists of three main steps: cropping a region of interest from a mosaic across all available filters; stacking the images to construct the data cube; and performing PSF matching on the cube. This package is intended for preparing data for studies involving multi-band photometry. The code is released under an MIT license and is available on GitHub together with a Jupyter tutorial notebook. The version used for this publication (v0.2.1) is archived on Zenodo.

astro-ph.IM

SAGUI: SED-based Segmentation of Multi-band Galaxy Images -- Application to JADES in GOODS-South

We present sagui, a modular framework for the analysis of multi-band imaging data in spatially resolved galaxies, with synergies to integral-field spectroscopy (IFS). Building on the spectro-spatial paradigm introduced by capivara for IFS data, sagui extends this approach to imaging datasets, enabling a coherent, pixel-level treatment of spatial and spectral information across multiple bands. The method follows a two-stage strategy: a starlet-based decomposition is first used to identify and mask spatial structures across multiple scales while suppressing noise, and a spectral-similarity analysis then partitions the image into coherent pixel groups that preserve spectral consistency. In addition to compact and high-contrast structures, the framework incorporates a dedicated statistical treatment, based on a copula transform, to identify and recover faint, diffuse low-surface-brightness components. We demonstrate the method across a diverse range of galaxy morphologies, highlighting its ability to characterize complex spatial structures, including clumps, bars, interacting systems, and low-surface-brightness features. As a case study, we apply it to eleven morphologically diverse galaxies from the James Webb Space Telescope Advanced Deep Extragalactic Survey in the GOODS--South field. sagui is released under an MIT license and is available at https://rafaelsdesouza.github.io/sagui/.

astro-ph.IM

Semi-Supervised Learning for Lensed Quasar Detection

Lensed quasars are key to many areas of study in astronomy, offering a unique probe into the intermediate and far universe. However, finding lensed quasars has proved difficult despite significant efforts from large collaborations. These challenges have limited catalogues of confirmed lensed quasars to the hundreds, despite theoretical predictions that they should be many times more numerous. We train machine learning classifiers to discover lensed quasar candidates. By using semi-supervised learning techniques we leverage the large number of potential candidates as unlabelled training data alongside the small number of known objects, greatly improving model performance. We present our two most successful models: (1) a variational autoencoder trained on millions of quasars to reduce the dimensionality of images for input to a dense neural network classifier that can make accurate predictions and (2) a convolutional neural network trained on a mix of labelled and unlabelled data via virtual adversarial training. These models are both capable of producing high-quality candidates, as evidenced by our discovery of GRALJ140833.73+042229.98. The success of our classifier, which uses only multi-band images, is particularly exciting as it can be combined with existing classifiers, which use other data than images, to improve the classifications of both models and discover more lensed quasars.

astro-ph.IM

AT2022zod: An Unusual Tidal Disruption Event in an Elliptical Galaxy at Redshift 0.11

Tidal Disruption Events (TDEs) have long been hypothesized as valuable indicators of black holes, offering insight into their demographics and behaviour out to high redshift. TDEs have also enabled the discovery of a few Massive Black Holes (MBHs) with inferred masses of $10^4$--$10^6\,M_\odot$, often associated with the nuclei of dwarf galaxies or ultra-compact dwarf galaxies (UCDs). Here we present AT2022zod, an extreme and short-lived optical flare in an elliptical galaxy at z=0.11, residing within 3kpc of the galaxy's centre. Its luminosity and ~30-day duration make it unlikely to have originated from the host galaxy's central supermassive black hole (SMBH), which estimate to have a mass around $10^8\,M_\odot$. Assuming that the emission mechanism is consistent with known observed TDEs, we find that such a rapidly evolving transient could either be produced by a MBH in the intermediate-mass range or, alternatively, result from the tidal disruption of a star on a non-parabolic orbit around the central SMBH. We suggest that the most plausible origin for AT2022zod is the tidal disruption of a star by a MBH embedded in a UCD. As the Vera C. Rubin Observatory's Legacy Survey of Space and Time comes online, we propose that AT2022zod serves as an important event for refining search strategies and characterization techniques for intermediate-mass black holes.

astro-ph.HE

A Brief History of Inference in Astronomy

In this short review, we trace the evolution of inference in astronomy, highlighting key milestones rather than providing an exhaustive survey. We focus on the shift from classical optimization to Bayesian inference, the rise of gradient-based methods fueled by advances in deep learning, and the emergence of adaptive models that shape the very design of scientific datasets. Understanding this shift is essential for appreciating the current landscape of astronomical research and the future it is helping to build.

astro-ph.IM

Memorization: A Close Look at Books

To what extent can entire books be extracted from LLMs? Using the Llama 3 70B family of models, and the "prefix-prompting" extraction technique, we were able to auto-regressively reconstruct, with a very high level of similarity, one entire book (Alice's Adventures in Wonderland) from just the first 500 tokens. We were also able to obtain high extraction rates on several other books, piece-wise. However, these successes do not extend uniformly to all books. We show that extraction rates of books correlate with book popularity and thus, likely duplication in the training data. We also confirm the undoing of mitigations in the instruction-tuned Llama 3.1, following recent work (Nasr et al., 2025). We further find that this undoing comes from changes to only a tiny fraction of weights concentrated primarily in the lower transformer blocks. Our results provide evidence of the limits of current regurgitation mitigation strategies and introduce a framework for studying how fine-tuning affects the retrieval of verbatim memorization in aligned LLMs.

cs.CL

Gaia GraL: Gaia gravitational lens systems IX. Using XGBoost to explore the Gaia Focused Product Release GravLens catalogue

Aims. Quasar strong gravitational lenses are important tools for putting constraints on the dark matter distribution, dark energy contribution, and the Hubble-Lemaitre parameter. We aim to present a new supervised machine learning-based method to identify these lenses in large astrometric surveys. The Gaia Focused Product Release (FPR) GravLens catalogue is designed for the identification of multiply imaged quasars, as it provides astrometry and photometry of all sources in the field of 4.7 million quasars. Methods. Our new approach for automatically identifying four-image lens configurations in large catalogues is based on the eXtreme Gradient Boosting classification algorithm. To train this supervised algorithm, we performed realistic simulations of lenses with four images that account for the statistical distribution of the morphology of the deflecting halos as measured in the EAGLE simulation. We identified the parameters discriminant for the classification and performed two different trainings, namely, with and without distance information. Results. The performances of this method on the simulated data are quite good, with a true positive rate and a true negative rate of about 99.99% and 99.84%, respectively. Our validation of the method on a small set of known quasar lenses demonstrates its efficiency, with 75% of known lenses being correctly identified. We applied our algorithm (both trainings) to more than 0.9 million quadruplets selected from the Gaia FPR GravLens catalogue. We derived a list of 1127 candidates with at least one score larger than 0.75, where each candidate has two scores -- one from the model trained with distance information and one from the model trained without distance information -- and including 201 very good candidates with both high scores.

astro-ph.GA

From Galaxy Zoo DECaLS to BASS/MzLS: detailed galaxy morphology classification with unsupervised domain adaption

The DESI Legacy Imaging Surveys (DESI-LIS) comprise three distinct surveys: the Dark Energy Camera Legacy Survey (DECaLS), the Beijing-Arizona Sky Survey (BASS), and the Mayall z-band Legacy Survey (MzLS). The citizen science project Galaxy Zoo DECaLS 5 (GZD-5) has provided extensive and detailed morphology labels for a sample of 253,287 galaxies within the DECaLS survey. This dataset has been foundational for numerous deep learning-based galaxy morphology classification studies. However, due to differences in signal-to-noise ratios and resolutions between the DECaLS images and those from BASS and MzLS (collectively referred to as BMz), a neural network trained on DECaLS images cannot be directly applied to BMz images due to distributional mismatch. In this study, we explore an unsupervised domain adaptation (UDA) method that fine-tunes a source domain model trained on DECaLS images with GZD-5 labels to BMz images, aiming to reduce bias in galaxy morphology classification within the BMz survey. Our source domain model, used as a starting point for UDA, achieves performance on the DECaLS galaxies' validation set comparable to the results of related works. For BMz galaxies, the fine-tuned target domain model significantly improves performance compared to the direct application of the source domain model, reaching a level comparable to that of the source domain. We also release a catalogue of detailed morphology classifications for 248,088 galaxies within the BMz survey, accompanied by usage recommendations.

astro-ph.GA

Observing the Galactic Underworld: Predicting photometry and astrometry from compact remnant microlensing events

Isolated black holes (BHs) and neutron stars (NSs) are largely undetectable across the electromagnetic spectrum. For this reason, our only real prospect of observing these isolated compact remnants is via microlensing; a feat recently performed for the first time. However, characterisation of the microlensing events caused by BHs and NSs is still in its infancy. In this work, we perform N-body simulations to explore the frequency and physical characteristics of microlensing events across the entire sky. Our simulations find that every year we can expect $88_{-6}^{+6}$ BH, $6.8_{-1.6}^{+1.7}$ NS and $20^{+30}_{-20}$ stellar microlensing events which cause an astrometric shift larger than 2~mas. Similarly, we can expect $21_{-3}^{+3}$ BH, $18_{-3}^{+3}$ NS and $7500_{-500}^{+500}$ stellar microlensing events which cause a bump magnitude larger than 1~mag. Leveraging a more comprehensive dynamical model than prior work, we predict the fraction of microlensing events caused by BHs as a function of Einstein time to be smaller than previously thought. Comparison of our microlensing simulations to events in Gaia finds good agreement. Finally, we predict that in the combination of Gaia and GaiaNIR data there will be $14700_{-900}^{+600}$ BH and $1600_{-200}^{+300}$ NS events creating a centroid shift larger than 1~mas and $330_{-120}^{+100}$ BH and $310_{-100}^{+110}$ NS events causing bump magnitudes $> 1$. Of these, $<10$ BH and $5_{-5}^{+10}$ NS events should be detectable using current analysis techniques. These results inform future astrometric mission design, such as GaiaNIR, as they indicate that, compared to stellar events, there are fewer observable BH events than previously thought.

astro-ph.GA

AstroInformatics: Recommendations for Global Cooperation

Policy Brief on "AstroInformatics, Recommendations for Global Collaboration", distilled from panel discussions during S20 Policy Webinar on Astroinformatics for Sustainable Development held on 6-7 July 2023. The deliberations encompassed a wide array of topics, including broad astroinformatics, sky surveys, large-scale international initiatives, global data repositories, space-related data, regional and international collaborative efforts, as well as workforce development within the field. These discussions comprehensively addressed the current status, notable achievements, and the manifold challenges that the field of astroinformatics currently confronts. The G20 nations present a unique opportunity due to their abundant human and technological capabilities, coupled with their widespread geographical representation. Leveraging these strengths, significant strides can be made in various domains. These include, but are not limited to, the advancement of STEM education and workforce development, the promotion of equitable resource utilization, and contributions to fields such as Earth Science and Climate Science. We present a concise overview, followed by specific recommendations that pertain to both ground-based and space data initiatives. Our team remains readily available to furnish further elaboration on any of these proposals as required. Furthermore, we anticipate further engagement during the upcoming G20 presidencies in Brazil (2024) and South Africa (2025) to ensure the continued discussion and realization of these objectives. The policy webinar took place during the G20 presidency in India (2023). Notes based on the seven panels will be separately published.

astro-ph.IM

GraL spectroscopic identification of multiply imaged quasars

Gravitational lensing is proven to be one of the most efficient tools for studying the Universe. The spectral confirmation of such sources requires extensive calibration. This paper discusses the spectral extraction technique for the case of multiple source spectra being very near each other. Using the masking technique, we first detect high Signal-to-Noise (S/N) peaks in the CCD spectral image corresponding to the location of the source spectra. This technique computes the cumulative signal using a weighted sum, yielding a reliable approximation for the total counts contributed by each source spectrum. We then proceed with the subtraction of the contaminating spectra. Applying this method, we confirm the nature of 11 lensed quasar candidates.

astro-ph.CO

Gaia GraL: Gaia DR2 Gravitational Lens Systems. VIII. A radio census of lensed systems

We present radio observations of 24 confirmed and candidate strongly lensed quasars identified by the Gaia Gravitational Lenses (GraL) working group. We detect radio emission from 8 systems in 5.5 and 9 GHz observations with the Australia Telescope Compact Array (ATCA), and 12 systems in 6 GHz observations with the Karl G. Jansky Very Large Array (VLA). The resolution of our ATCA observations is insufficient to resolve the radio emission into multiple lensed images, but we do detect multiple images from 11 VLA targets. We have analysed these systems using our observations in conjunction with existing optical measurements, including measuring offsets between the radio and optical positions, for each image and building updated lens models. These observations significantly expand the existing sample of lensed radio quasars, suggest that most lensed systems are detectable at radio wavelengths with targeted observations, and demonstrate the feasibility of population studies with high resolution radio imaging.

astro-ph.GA

From Images to Features: Unbiased Morphology Classification via Variational Auto-Encoders and Domain Adaptation

We present a novel approach for the dimensionality reduction of galaxy images by leveraging a combination of variational auto-encoders (VAE) and domain adaptation (DA). We demonstrate the effectiveness of this approach using a sample of low redshift galaxies with detailed morphological type labels from the Galaxy-Zoo DECaLS project. We show that 40-dimensional latent variables can effectively reproduce most morphological features in galaxy images. To further validate the effectiveness of our approach, we utilised a classical random forest (RF) classifier on the 40-dimensional latent variables to make detailed morphology feature classifications. This approach performs similarly to a direct neural network application on galaxy images. We further enhance our model by tuning the VAE network via DA using galaxies in the overlapping footprint of DECaLS and BASS+MzLS, enabling the unbiased application of our model to galaxy images in both surveys. We observed that DA led to even better morphological feature extraction and classification performance. Overall, this combination of VAE and DA can be applied to achieve image dimensionality reduction, defect image identification, and morphology classification in large optical surveys.

astro-ph.GA

Detection of open cluster rotation fields from Gaia EDR3 proper motions

Context. Most stars from in groups which with time disperse, building the field population of their host galaxy. In the Milky Way, open clusters have been continuously forming in the disk up to the present time, providing it with stars spanning a broad range of ages and masses. Observations of the details of cluster dissolution are, however, scarce. One of the main difficulties is obtaining a detailed characterisation of the internal cluster kinematics, which requires very high quality proper motions. For open clusters, which are typically loose groups with some tens to hundreds of members, there is the additional difficulty of inferring kinematic structures from sparse and irregular distributions of stars. Aims. Here, we aim to analyse internal stellar kinematics of open clusters, and identify rotation, expansion or contraction patterns. Methods. We use Gaia Early Data Release 3 (EDR3) astrometry and Integrated Nested Laplace Approximations to perform vector-field inference and create spatio-kinematic maps of 1237 open clusters. The sample is composed of clusters for which individual stellar memberships were known, thus minimising contamination from field stars in the velocity maps. Projection effects were corrected using EDR3 data complemented with radial velocities from Gaia Data Release 2 and other surveys. Results. We report the detection of rotation patterns in 8 open clusters. Nine additional clusters display possible rotation signs. We also observe 14 expanding clusters, with 15 other objects showing possible expansion patterns. Contraction is evident in two clusters, with one additional cluster presenting a more uncertain detection. In total, 53 clusters are found to display kinematic structures. Within these, elongated spatial distributions suggesting tidal tails are found in 5 clusters. [abridged]

astro-ph.GA

The spin axes of globular clusters and correlations with gamma-ray emission

A growing number of Milky Way globular clusters have been identified to possess a noticeable degree of solid-body rotation. For several clusters, the combination of stellar proper motions and radial velocities allows for 3-dimensional spin axes to be extracted. In this paper we consider the orientations of these spin axes, and ask whether they are correlated with any other properties of the clusters -- either global properties to do with their orbits and origin, or internal properties related to the cluster composition. We discuss the possibility of alignments between the spin axes of globular clusters, chemodynamical groupings, and their orbital poles. We also point out a previously unidentified negative correlation between the measured gamma-ray emissivities and the inclination of the globular cluster spins with respect to the line of sight. Given that this correlation is not present in other wavelengths, we cannot conclusively attribute it solely to sampling bias. If the correlation holds up to scrutiny with more data, it may be indicative of sources of anisotropic gamma-ray emission in globular clusters. We discuss the plausibility of such an anisotropy arising from a population of dynamically formed millisecond pulsars with some degree of spin-orbit alignment.

astro-ph.GA

A graph-based spectral classification of Type II supernovae

Given the ever-increasing number of time-domain astronomical surveys, employing robust, interpretative, and automated data-driven classification schemes is pivotal. Based on graph theory, we present new data-driven classification heuristics for spectral data. A spectral classification scheme of Type II supernovae (SNe II) is proposed based on the phase relative to the maximum light in the $V$ band and the end of the plateau phase. We utilize a compiled optical data set that comprises 145 SNe and 1595 optical spectra in 4000-9000 $\overset{\circ}{\mathrm {A}}$. Our classification method naturally identifies outliers and arranges the different SNe in terms of their major spectral features. We compare our approach to the off-the-shelf umap manifold learning and show that both strategies are consistent with a continuous variation of spectral types rather than discrete families. The automated classification naturally reflects the fast evolution of Type II SNe around the maximum light while showcasing their homogeneity close to the end of the plateau phase. The scheme we develop could be more widely applicable to unsupervised time series classification or characterisation of other functional data.

astro-ph.IM

Are classification metrics good proxies for SN Ia cosmological constraining power?

Context: When selecting a classifier to use for a supernova Ia (SN Ia) cosmological analysis, it is common to make decisions based on metrics of classification performance, i.e. contamination within the photometrically classified SN Ia sample, rather than a measure of cosmological constraining power. If the former is an appropriate proxy for the latter, this practice would save those designing an analysis pipeline from the computational expense of a full cosmology forecast. Aims: This study tests the assumption that classification metrics are an appropriate proxy for cosmology metrics. Methods: We emulate photometric SN Ia cosmology samples with controlled contamination rates of individual contaminant classes and evaluate each of them under a set of classification metrics. We then derive cosmological parameter constraints from all samples under two common analysis approaches and quantify the impact of contamination by each contaminant class on the resulting cosmological parameter estimates. Results: We observe that cosmology metrics are sensitive to both the contamination rate and the class of the contaminating population, whereas the classification metrics are insensitive to the latter. Conclusions: We therefore discourage exclusive reliance on classification-based metrics for cosmological analysis design decisions, e.g. classifier choice, and instead recommend optimizing using a metric of cosmological parameter constraining power.

astro-ph.CO

Fast emulation of cosmological density fields based on dimensionality reduction and supervised machine-learning

N-body simulations are the most powerful method to study the non-linear evolution of large-scale structure. However, they require large amounts of computational resources, making unfeasible their direct adoption in scenarios that require broad explorations of parameter spaces. In this work, we show that it is possible to perform fast dark matter density field emulations with competitive accuracy using simple machine-learning approaches. We build an emulator based on dimensionality reduction and machine learning regression combining simple Principal Component Analysis and supervised learning methods. For the estimations with a single free parameter, we train on the dark matter density parameter, $Ω_m$, while for emulations with two free parameters, we train on a range of $Ω_m$ and redshift. The method first adopts a projection of a grid of simulations on a given basis; then, a machine learning regression is trained on this projected grid. Finally, new density cubes for different cosmological parameters can be estimated without relying directly on new N-body simulations by predicting and de-projecting the basis coefficients. We show that the proposed emulator can generate density cubes at non-linear cosmological scales with density distributions within a few percent compared to the corresponding N-body simulations. The method enables gains of three orders of magnitude in CPU run times compared to performing a full N-body simulation while reproducing the power spectrum and bispectrum within $\sim 1\%$ and $\sim 3\%$, respectively, for the single free parameter emulation and $\sim 5\%$ and $\sim 15\%$ for two free parameters. This can significantly accelerate the generation of density cubes for a wide variety of cosmological models, opening the doors to previously unfeasible applications, such as parameter and model inferences at full survey scales as the ESA/NASA Euclid mission.

astro-ph.CO