SearcharxivSearch

arXiv subjects

Guillermo Cabrera-Vives

Publications and source records attributed to Guillermo Cabrera-Vives.

At least 19 recordsLinked to original sources

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model

We explore how an uncertainty-aware transformer-based architecture can leverage information embedded across the entire observed optical spectra of AGN, focusing on the algorithm's ability to predict unseen or masked parts of luminous AGN spectra. This provides a direct probe of the learnable correlations between AGN continua and broad lines. We introduce AGNFormer, a transformer model trained to predict the mean expected flux and variance in masked spectral regions (major broad lines to ${\pm}10^{4}$kms$^{-1}$; missing halves), inputting rest-frame spectral fluxes and uncertainties across the entire redshift range of the SDSS DR16 Quasar Catalogue. We evaluate the performance of the model on both full (no S/N limit) and high-quality (S/N > 10) spectral samples using the negative-log likelihood, and via comparisons with existing C IV and ly-a reconstruction algorithms. The model successfully reconstructs unseen AGN broad lines to better than 10-16% (4-8%) of the flux for the full (S/N > 10) test sets, up to an error floor of $\approx$2-6% of the flux at S/N $\approx$ 40, while predictions for larger unseen halves grow to 12-25% (5-15%) of the flux the further away they are from the cut-off wavelength of the seen input spectrum. Predictions faithfully reproduce the broad AGN spectral diversity across the entire optical and UV QSO main sequence parameter spaces, including both Gaussian and Lorentzian profile regimes, Feii complexes, and narrow emission lines. Performance is similar or better compared to previous spectral reconstruction algorithms. The high precision of the broad-line region reconstruction demonstrates that the method successfully aggregates information across the spectrum and highlights how the AGN continuum and weaker lines/complexes have the potential to assist astronomers in the extraction of the entire wealth of information embedded in AGN spectra.

astro-ph.GA

Beyond the Final Label: Exploiting the Untapped Potential of Classification Histories in Astronomical Light Curve Analysis

The Legacy Survey of Space and Time (LSST) on the Vera C. Rubin Observatory will generate a massive collection of time series (light curves) of the measured flux of transient and variable astronomical objects. With each new flux observation, light curve classifiers need to generate updated probability distributions over candidate classes, which will then be shared with the global community for the purpose of identifying interesting targets for follow-up observations as well as less time-sensitive analysis applications. Using the synthetic light curves and classification results of participating classifiers from the Extended LSST Astronomical Time-series Classification Challenge (ELAsTiCC), we investigate a novel framework to enhance existing light curve classifications by incorporating their classification histories and the temporal evolution of these histories. To demonstrate the potential of this approach, we introduce a model that combines a recurrent neural network and an additive attention module, which shows improved classification accuracy and more balanced precision-recall performance compared to existing classifiers from the challenge. Furthermore, at this stage, most, if not all, of the existing classifiers are evaluated by their final classification results on complete light curves; we propose new metrics that evaluate the stability, accuracy, and early classification performance of a classifier's predictions when using limited data by considering the Wasserstein distance between the temporally evolving classification probability distributions. Our metrics offer a more comprehensive perspective for model assessment by supplementing classical methods such as the confusion matrix and precision-recall.

astro-ph.IM

Caught in the web: galaxy mergers along cosmic filaments

Galaxy clusters grow through the accretion of galaxies from groups, filaments, and other clusters. During this process, galaxies may undergo pre-processing in lower-density environments, where galaxy-galaxy mergers and other interactions can significantly alter their properties prior to cluster infall. We investigate the role of galaxy mergers in the pre-processing of galaxies prior to cluster infall by studying the spatial distribution of mergers across the cosmic web. We use a sample of 43,922 galaxies targeted by the 4MOST CHANCES survey in and around 33 low-redshift clusters (z < 0.07). Using Zoobot, a deep-learning framework trained on Galaxy Zoo data, we identify 698 galaxy mergers. We measure their distances to cosmic web filaments and compare them with those of non-merging galaxies. We find that galaxy mergers are significantly closer to filaments than the non-merging galaxy population, with this trend being strongest beyond the cluster virial radius. This suggests that filaments provide conditions conducive to mergers, possibly moderating relative velocities and enhancing gas availability. Our findings support a scenario in which filaments play a key role in transforming galaxies through pre-processing by promoting mergers before they enter cluster cores where star formation quenches.

astro-ph.GA

Astromer 2

Foundational models have emerged as a powerful paradigm in deep learning field, leveraging their capacity to learn robust representations from large-scale datasets and effectively to diverse downstream applications such as classification. In this paper, we present Astromer 2 a foundational model specifically designed for extracting light curve embeddings. We introduce Astromer 2 as an enhanced iteration of our self-supervised model for light curve analysis. This paper highlights the advantages of its pre-trained embeddings, compares its performance with that of its predecessor, Astromer 1, and provides a detailed empirical analysis of its capabilities, offering deeper insights into the model's representations. Astromer 2 is pretrained on 1.5 million single-band light curves from the MACHO survey using a self-supervised learning task that predicts randomly masked observations within sequences. Fine-tuning on a smaller labeled dataset allows us to assess its performance in classification tasks. The quality of the embeddings is measured by the F1 score of an MLP classifier trained on Astromer-generated embeddings. Our results demonstrate that Astromer 2 significantly outperforms Astromer 1 across all evaluated scenarios, including limited datasets of 20, 100, and 500 samples per class. The use of weighted per-sample embeddings, which integrate intermediate representations from Astromer's attention blocks, is particularly impactful. Notably, Astromer 2 achieves a 15% improvement in F1 score on the ATLAS dataset compared to prior models, showcasing robust generalization to new datasets. This enhanced performance, especially with minimal labeled data, underscores the potential of Astromer 2 for more efficient and scalable light curve analysis.

astro-ph.IM

Leveraging pre-trained vision Transformers for multi-band photometric light curve classification

This study investigates the potential of a pre-trained vision Transformer (VT) model, specifically the Swin Transformer V2 (SwinV2), to classify photometric light curves without the need for feature extraction or multi-band preprocessing. The goal is to assess whether this image-based approach can accurately differentiate astronomical phenomena and serve as a viable option for working with multi-band photometric light curves. We transformed each multi-band light curve into an image. These images serve as input to the SwinV2 model, which is pre-trained on ImageNet-21K. The datasets employed include the public Catalog of Variable Stars from the Massive Compact Halo Object (MACHO) survey, using both one and two bands, and the first round of the recent Extended LSST Astronomical Time-Series Classification Challenge (ELAsTiCC), which includes six bands. The performance of the model was evaluated on six classes for the MACHO dataset and 20 distinct classes of variable stars and transient events for the ELAsTiCC dataset. The fine-tuned SwinV2 achieved better performance than models specifically designed for light curves, such as Astromer and the Astronomical Transformer for Time Series and Tabular Data (ATAT). When trained on the full MACHO dataset, it attained a macro F1-score of 80.2 and outperformed Astromer in single-band experiments. Incorporating a second band further improved performance, increasing the F1-score to 84.1. In the ELAsTiCC dataset, SwinV2 achieved a macro F1-score of 65.5, slightly surpassing ATAT by 1.3.

astro-ph.IM

ASTROCO: Self-Supervised Conformer-Style Transformers for Light-Curve Embeddings

We present AstroCo, a Conformer-style encoder for irregular stellar light curves. By combining attention with depthwise convolutions and gating, AstroCo captures both global dependencies and local features. On MACHO R-band, AstroCo outperforms Astromer v1 and v2, yielding 70 percent and 61 percent lower error respectively and a relative macro-F1 gain of about 7 percent, while producing embeddings that transfer effectively to few-shot classification. These results highlight AstroCo's potential as a strong and label-efficient foundation for time-domain astronomy.

astro-ph.IM

Astro-MoE: Mixture of Experts for Multiband Astronomical Time Series

Multiband astronomical time series exhibit heterogeneous variability patterns, sampling cadences, and signal characteristics across bands. Standard transformers apply shared parameters to all bands, potentially limiting their ability to model this rich structure. In this work, we introduce Astro-MoE, a foundational transformer architecture that enables dynamic processing via a Mixture of Experts module. We validate our model on both simulated (ELAsTiCC-1) and real-world datasets (Pan-STARRS1).

astro-ph.IM

Image-Based Multi-Survey Classification of Light Curves with a Pre-Trained Vision Transformer

We explore the use of Swin Transformer V2, a pre-trained vision Transformer, for photometric classification in a multi-survey setting by leveraging light curves from the Zwicky Transient Facility (ZTF) and the Asteroid Terrestrial-impact Last Alert System (ATLAS). We evaluate different strategies for integrating data from these surveys and find that a multi-survey architecture which processes them jointly achieves the best performance. These results highlight the importance of modeling survey-specific characteristics and cross-survey interactions, and provide guidance for building scalable classifiers for future time-domain astronomy.

astro-ph.IM

Applying Vision Transformers on Spectral Analysis of Astronomical Objects

We apply pre-trained Vision Transformers (ViTs), originally developed for image recognition, to the analysis of astronomical spectral data. By converting traditional one-dimensional spectra into two-dimensional image representations, we enable ViTs to capture both local and global spectral features through spatial self-attention. We fine-tune a ViT pretrained on ImageNet using millions of spectra from the SDSS and LAMOST surveys, represented as spectral plots. Our model is evaluated on key tasks including stellar object classification and redshift ($z$) estimation, where it demonstrates strong performance and scalability. We achieve classification accuracy higher than Support Vector Machines and Random Forests, and attain $R^2$ values comparable to AstroCLIP's spectrum encoder, even when generalizing across diverse object types. These results demonstrate the effectiveness of using pretrained vision models for spectroscopic data analysis. To our knowledge, this is the first application of ViTs to large-scale, which also leverages real spectroscopic data and does not rely on synthetic inputs.

astro-ph.IM

Uncertainty estimation for time series classification: Exploring predictive uncertainty in transformer-based models for variable stars

Classifying variable stars is key for understanding stellar evolution and galactic dynamics. With the demands of large astronomical surveys, machine learning models, especially attention-based neural networks, have become the state-of-the-art. While achieving high accuracy is crucial, enhancing model interpretability and uncertainty estimation is equally important to ensure that insights are both reliable and comprehensible. We aim to enhance transformer-based models for classifying astronomical light curves by incorporating uncertainty estimation techniques to detect misclassified instances. We tested our methods on labeled datasets from MACHO, OGLE-III, and ATLAS, introducing a framework that significantly improves the reliability of automated classification for the next-generation surveys. We used Astromer, a transformer-based encoder designed for capturing representations of single-band light curves. We enhanced its capabilities by applying three methods for quantifying uncertainty: Monte Carlo Dropout (MC Dropout), Hierarchical Stochastic Attention (HSA), and a novel hybrid method combining both approaches, which we have named Hierarchical Attention with Monte Carlo Dropout (HA-MC Dropout). We compared these methods against a baseline of deep ensembles (DEs). To estimate uncertainty estimation scores for the misclassification task, we selected Sampled Maximum Probability (SMP), Probability Variance (PV), and Bayesian Active Learning by Disagreement (BALD) as uncertainty estimates. In predictive performance tests, HA-MC Dropout outperforms the baseline, achieving macro F1-scores of 79.8+-0.5 on OGLE, 84+-1.3 on ATLAS, and 76.6+-1.8 on MACHO. When comparing the PV score values, the quality of uncertainty estimation by HA-MC Dropout surpasses that of all other methods, with improvements of 2.5+-2.3 for MACHO, 3.3+-2.1 for ATLAS and 8.5+-1.6 for OGLE-III.

astro-ph.IM

Transformer-Based Astronomical Time Series Model with Uncertainty Estimation for Detecting Misclassified Instances

In this work, we present a framework for estimating and evaluating uncertainty in deep-attention-based classifiers for light curves for variable stars. We implemented three techniques, Deep Ensembles (DEs), Monte Carlo Dropout (MCD) and Hierarchical Stochastic Attention (HSA) and evaluated models trained on three astronomical surveys. Our results demonstrate that MCD and HSA offers a competitive and computationally less expensive alternative to DE, allowing the training of transformers with the ability to estimate uncertainties for large-scale light curve datasets. We conclude that the quality of the uncertainty estimation is evaluated using the ROC AUC metric.

astro-ph.IM

A Novel Optimal Transport-Based Approach for Interpolating Spectral Time Series: Paving the Way for Photometric Classification of Supernovae

This paper introduces a novel method for creating spectral time series, which can be used for generating synthetic light curves for photometric classification but also for applications like K-corrections and bolometric corrections. This approach is particularly valuable in the era of large astronomical surveys, where it can significantly enhance the analysis and understanding of an increasing number of SNe, even in the absence of extensive spectroscopic data. methods: By employing interpolations based on optimal transport theory, starting from a spectroscopic sequence, we derive weighted average spectra with high cadence. The weights incorporate an uncertainty factor for penalizing interpolations between spectra that show significant epoch differences and lead to a poor match between the synthetic and observed photometry. results: Our analysis reveals that even with phase difference of up to 40 days between pairs of spectra, optical transport can generate interpolated spectral time series that closely resemble the original ones. Synthetic photometry extracted from these spectral time series aligns well with observed photometry. The best results are achieved in the V band, with relative residuals of less than 10% for 87% and 84% of the data for type Ia and II, respectively. For the B, g, R and r bands, the relative residuals are between 65% and 87% within the previously mentioned 10% threshold for both classes. The worse results correspond to the i and I bands where, in the case, of SN~Ia the values drop to 53% and 42%, respectively. conclusions: We introduce a new method for constructing spectral time series for individual SNe starting from a sparse spectroscopic sequence, and demonstrate its capability to produce reliable light curves that can be used for photometric classification.

astro-ph.HE

Mitigating Bias in Deep Learning: Training Unbiased Models on Biased Data for the Morphological Classification of Galaxies

Galaxy morphologies and their relation with physical properties have been a relevant subject of study in the past. Most galaxy morphology catalogs have been labelled by human annotators or by machine learning models trained on human labelled data. Human generated labels have been shown to contain biases in terms of the observational properties of the data, such as image resolution. These biases are independent of the annotators, that is, are present even in catalogs labelled by experts. In this work, we demonstrate that training deep learning models on biased galaxy data produce biased models, meaning that the biases in the training data are transferred to the predictions of the new models. We also propose a method to train deep learning models that considers this inherent labelling bias, to obtain a de-biased model even when training on biased data. We show that models trained using our deep de-biasing method are capable of reducing the bias of human labelled datasets.

astro-ph.GA

Domain Adaptation via Minimax Entropy for Real/Bogus Classification of Astronomical Alerts

Time domain astronomy is advancing towards the analysis of multiple massive datasets in real time, prompting the development of multi-stream machine learning models. In this work, we study Domain Adaptation (DA) for real/bogus classification of astronomical alerts using four different datasets: HiTS, DES, ATLAS, and ZTF. We study the domain shift between these datasets, and improve a naive deep learning classification model by using a fine tuning approach and semi-supervised deep DA via Minimax Entropy (MME). We compare the balanced accuracy of these models for different source-target scenarios. We find that both the fine tuning and MME models improve significantly the base model with as few as one labeled item per class coming from the target dataset, but that the MME does not compromise its performance on the source dataset.

astro-ph.IM

Con$^{2}$DA: Simplifying Semi-supervised Domain Adaptation by Learning Consistent and Contrastive Feature Representations

In this work, we present Con$^{2}$DA, a simple framework that extends recent advances in semi-supervised learning to the semi-supervised domain adaptation (SSDA) problem. Our framework generates pairs of associated samples by performing stochastic data transformations to a given input. Associated data pairs are mapped to a feature representation space using a feature extractor. We use different loss functions to enforce consistency between the feature representations of associated data pairs of samples. We show that these learned representations are useful to deal with differences in data distributions in the domain adaptation problem. We performed experiments to study the main components of our model and we show that (i) learning of the consistent and contrastive feature representations is crucial to extract good discriminative features across different domains, and ii) our model benefits from the use of strong augmentation policies. With these findings, our method achieves state-of-the-art performances in three benchmark datasets for SSDA.

cs.LG

Positional Encodings for Light Curve Transformers: Playing with Positions and Attention

We conducted empirical experiments to assess the transferability of a light curve transformer to datasets with different cadences and magnitude distributions using various positional encodings (PEs). We proposed a new approach to incorporate the temporal information directly to the output of the last attention layer. Our results indicated that using trainable PEs lead to significant improvements in the transformer performances and training times. Our proposed PE on attention can be trained faster than the traditional non-trainable PE transformer while achieving competitive results when transfered to other datasets.

astro-ph.IM

Multi-Class Deep SVDD: Anomaly Detection Approach in Astronomy with Distinct Inlier Categories

With the increasing volume of astronomical data generated by modern survey telescopes, automated pipelines and machine learning techniques have become crucial for analyzing and extracting knowledge from these datasets. Anomaly detection, i.e. the task of identifying irregular or unexpected patterns in the data, is a complex challenge in astronomy. In this paper, we propose Multi-Class Deep Support Vector Data Description (MCDSVDD), an extension of the state-of-the-art anomaly detection algorithm One-Class Deep SVDD, specifically designed to handle different inlier categories with distinct data distributions. MCDSVDD uses a neural network to map the data into hyperspheres, where each hypersphere represents a specific inlier category. The distance of each sample from the centers of these hyperspheres determines the anomaly score. We evaluate the effectiveness of MCDSVDD by comparing its performance with several anomaly detection algorithms on a large dataset of astronomical light-curves obtained from the Zwicky Transient Facility. Our results demonstrate the efficacy of MCDSVDD in detecting anomalous sources while leveraging the presence of different inlier categories. The code and the data needed to reproduce our results are publicly available at https://github.com/mperezcarrasco/AnomalyALeRCE.

cs.LG

Multi-scale stamps for real-time classification of alert streams

In recent years, automatic classifiers of image cutouts (also called "stamps") have shown to be key for fast supernova discovery. The Vera C. Rubin Observatory will distribute about ten million alerts with their respective stamps each night, enabling the discovery of approximately one million supernovae each year. A growing source of confusion for these classifiers is the presence of satellite glints, sequences of point-like sources produced by rotating satellites or debris. The currently planned Rubin stamps will have a size smaller than the typical separation between these point sources. Thus, a larger field of view stamp could enable the automatic identification of these sources. However, the distribution of larger stamps would be limited by network bandwidth restrictions. We evaluate the impact of using image stamps of different angular sizes and resolutions for the fast classification of events (AGNs, asteroids, bogus, satellites, SNe, and variable stars), using data from the Zwicky Transient Facility. We compare four scenarios: three with the same number of pixels (small field of view with high resolution, large field of view with low resolution, and a multi-scale proposal) and a scenario with the full stamp that has a larger field of view and higher resolution. Compared to small field of view stamps, our multi-scale strategy reduces misclassifications of satellites as asteroids or supernovae, performing on par with high-resolution stamps that are 15 times heavier. We encourage Rubin and its Science Collaborations to consider the benefits of implementing multi-scale stamps as a possible update to the alert specification.

astro-ph.IM