SearcharxivSearch

arXiv subjects

G. Cabrera-Vives

Publications and source records attributed to G. Cabrera-Vives.

15 recordsLinked to original sources

Querying an astronomical database using large language models: the ALeRCE text-to-SQL system

We develop a text-to-SQL (structured query language) system based on large language models (LLMs) using in-context learning and apply it to the Automatic Learning for the Rapid Classification of Events (ALeRCE) astronomical database. ALeRCE is a community broker for the Zwicky Transient Facility and the Vera C. Rubin Observatory. The system enables users to query the database in natural language (NL) and generates executable SQL queries. To develop and evaluate the system, we constructed a dataset of 110 NL/SQL pairs. We propose a step-by-step generation framework comprising four modules: schema linking, query classification, prompt decomposition, and self-correction. The performance of thirteen LLMs is evaluated using in-context learning and prompt engineering techniques. Text-to-SQL performance is assessed using the perfect-match (PM) rate for row identifiers (e.g., object identifiers) and column identifiers (i.e., column names). The proposed step-by-step framework consistently outperforms a direct-inference baseline, while the self-correction module consistently reduces execution errors. For Claude Opus 4.6, PM performance on row (column) identifiers is high for simple queries, reaching 0.97 (0.94), and decreases with query complexity to 0.44 (0.72) for medium queries and 0.59 (0.49) for hard queries. Among the thirteen evaluated models, the best-performing LLMs for the text-to-SQL task are Claude Opus 4.6, Gemini 2.5 Pro, Gemini 3 Flash, and GPT-5.2-Codex.

astro-ph.IM

ALeRCE light curve classifier: Tidal disruption event expansion pack

ALeRCE (Automatic Learning for the Rapid Classification of Events) processes the Zwicky Transient Facility (ZTF) alert stream in preparation for the Vera C. Rubin Observatory, classifying objects using a broad taxonomy. The ALeRCE light curve classifier is a balanced random forest (BRF) algorithm with a two-level scheme that uses variability features from the ZTF alert stream and colors from AllWISE and ZTF photometry. This work presents an updated version of the ALeRCE broker light curve classifier that includes tidal disruption events (TDEs) as a new subclass. We incorporated 24 new features, including the distance to the nearest source detected in ZTF science images and a parametric model of the power-law decay for transients. The labeled set was expanded to 219792 spectroscopically classified sources, including 60 TDEs. To integrate TDEs into ALeRCE's taxonomy, we identified specific characteristics that distinguish them from other transients: their central position in a galaxy, their typical decay pattern when fully disrupted, and their lack of color variability after disruption. We developed features to separate TDEs from other transient events. The updated classifier improves performance across all classes and integrates the TDE class with 91 percent recall. It also identifies a large number of potential TDE candidates in the ZTF alert stream's unlabeled data.

astro-ph.IM

ATAT: Astronomical Transformer for time series And Tabular data

The advent of next-generation survey instruments, such as the Vera C. Rubin Observatory and its Legacy Survey of Space and Time (LSST), is opening a window for new research in time-domain astronomy. The Extended LSST Astronomical Time-Series Classification Challenge (ELAsTiCC) was created to test the capacity of brokers to deal with a simulated LSST stream. We describe ATAT, the Astronomical Transformer for time series And Tabular data, a classification model conceived by the ALeRCE alert broker to classify light-curves from next-generation alert streams. ATAT was tested in production during the first round of the ELAsTiCC campaigns. ATAT consists of two Transformer models that encode light curves and features using novel time modulation and quantile feature tokenizer mechanisms, respectively. ATAT was trained on different combinations of light curves, metadata, and features calculated over the light curves. We compare ATAT against the current ALeRCE classifier, a Balanced Hierarchical Random Forest (BHRF) trained on human-engineered features derived from light curves and metadata. When trained on light curves and metadata, ATAT achieves a macro F1-score of 82.9 +- 0.4 in 20 classes, outperforming the BHRF model trained on 429 features, which achieves a macro F1-score of 79.4 +- 0.1. The use of Transformer multimodal architectures, combining light curves and tabular data, opens new possibilities for classifying alerts from a new generation of large etendue telescopes, such as the Vera C. Rubin Observatory, in real-world brokering scenarios.

astro-ph.IM

Astrometric and photometric characterization of $η$ Tel B combining two decades of observations

$η$ Tel is an 18 Myr system with a 2.09 M$_{\odot}$ A-type star and an M7-M8 brown dwarf companion, $η$ Tel B, separated by 4.2'' (208 au). High-contrast imaging campaigns over 20 years have enabled orbital and photometric characterization. $η$ Tel B, bright and on a wide orbit, is ideal for detailed examination. We analyzed three new SPHERE/IRDIS coronagraphic observations to explore $η$ Tel B's orbital parameters, contrast, and surroundings, aiming to detect a circumplanetary disk or close companion. Reduced IRDIS data achieved a contrast of 1.0$\times 10^{-5}$, enabling astrometric measurements with uncertainties of 4 mas in separation and 0.2 degrees in position angle, the smallest so far. With a contrast of 6.8 magnitudes in the H band, $η$ Tel B's separation and position angle were measured as 4.218'' and 167.3 degrees, respectively. Orbital analysis using Orvara code, considering Gaia-Hipparcos acceleration, revealed a low eccentric orbit (e $\sim$ 0.34), inclination of 81.9 degrees, and semi-major axis of 218 au. $η$ Tel B's mass was determined to be 48 \MJup, consistent with previous calculations. No significant residual indicating a satellite or disk around $η$ Tel B was detected. Detection limits ruled out massive objects around $η$ Tel B with masses down to 1.6 \MJup at a separation of 33 au.

astro-ph.SR

Persistent and occasional: searching for the variable population of the ZTF/4MOST sky using ZTF data release 11

We present a variability, color and morphology based classifier, designed to identify transients, persistently variable, and non-variable sources, from the Zwicky Transient Facility (ZTF) Data Release 11 (DR11) light curves of extended and point sources. The main motivation to develop this model was to identify active galactic nuclei (AGN) at different redshift ranges to be observed by the 4MOST ChANGES project. Still, it serves as a more general time-domain astronomy study. The model uses nine colors computed from CatWISE and PS1, a morphology score from PS1, and 61 single-band variability features computed from the ZTF DR11 g and r light curves. We trained two versions of the model, one for each ZTF band. We used a hierarchical local classifier per parent node approach, where each node was composed of a balanced random forest model. We adopted a 17-class taxonomy, including non-variable stars and galaxies, three transient classes, five classes of stochastic variables, and seven classes of periodic variables. The macro averaged precision, recall and F1-score are 0.61, 0.75, and 0.62 for the g-band model, and 0.60, 0.74, and 0.61, for the r-band model. When grouping the four AGN classes into one single class, its precision, recall, and F1-score are 1.00, 0.95, and 0.97, respectively, for both the g and r bands. We applied the model to all the sources in the ZTF/4MOST overlapping sky, avoiding ZTF fields covering the Galactic bulge, including 86,576,577 light curves in the g-band and 140,409,824 in the r-band. Only 0.73\% of the g-band light curves and 2.62\% of the r-band light curves were classified as stochastic, periodic, or transient with high probability ($P_{init}\geq0.9$). We found that, in general, more reliable results are obtained when using the g-band model. Using the latter, we identified 384,242 AGN candidates, 287,156 of which have $P_{init}\geq0.9$.

astro-ph.IM

ASTROMER: A transformer-based embedding for the representation of light curves

Taking inspiration from natural language embeddings, we present ASTROMER, a transformer-based model to create representations of light curves. ASTROMER was pre-trained in a self-supervised manner, requiring no human-labeled data. We used millions of R-band light sequences to adjust the ASTROMER weights. The learned representation can be easily adapted to other surveys by re-training ASTROMER on new sources. The power of ASTROMER consists of using the representation to extract light curve embeddings that can enhance the training of other models, such as classifiers or regressors. As an example, we used ASTROMER embeddings to train two neural-based classifiers that use labeled variable stars from MACHO, OGLE-III, and ATLAS. In all experiments, ASTROMER-based classifiers outperformed a baseline recurrent neural network trained on light curves directly when limited labeled data was available. Furthermore, using ASTROMER embeddings decreases computational resources needed while achieving state-of-the-art results. Finally, we provide a Python library that includes all the functionalities employed in this work. The library, main code, and pre-trained weights are available at https://github.com/astromer-science

astro-ph.IM

Searching for changing-state AGNs in massive datasets -- I: applying deep learning and anomaly detection techniques to find AGNs with anomalous variability behaviours

The classic classification scheme for Active Galactic Nuclei (AGNs) was recently challenged by the discovery of the so-called changing-state (changing-look) AGNs (CSAGNs). The physical mechanism behind this phenomenon is still a matter of open debate and the samples are too small and of serendipitous nature to provide robust answers. In order to tackle this problem, we need to design methods that are able to detect AGN right in the act of changing-state. Here we present an anomaly detection (AD) technique designed to identify AGN light curves with anomalous behaviors in massive datasets. The main aim of this technique is to identify CSAGN at different stages of the transition, but it can also be used for more general purposes, such as cleaning massive datasets for AGN variability analyses. We used light curves from the Zwicky Transient Facility data release 5 (ZTF DR5), containing a sample of 230,451 AGNs of different classes. The ZTF DR5 light curves were modeled with a Variational Recurrent Autoencoder (VRAE) architecture, that allowed us to obtain a set of attributes from the VRAE latent space that describes the general behaviour of our sample. These attributes were then used as features for an Isolation Forest (IF) algorithm, that is an anomaly detector for a "one class" kind of problem. We used the VRAE reconstruction errors and the IF anomaly score to select a sample of 8,809 anomalies. These anomalies are dominated by bogus candidates, but we were able to identify 75 promising CSAGN candidates.

astro-ph.IM

The effect of phased recurrent units in the classification of multiple catalogs of astronomical lightcurves

In the new era of very large telescopes, where data is crucial to expand scientific knowledge, we have witnessed many deep learning applications for the automatic classification of lightcurves. Recurrent neural networks (RNNs) are one of the models used for these applications, and the LSTM unit stands out for being an excellent choice for the representation of long time series. In general, RNNs assume observations at discrete times, which may not suit the irregular sampling of lightcurves. A traditional technique to address irregular sequences consists of adding the sampling time to the network's input, but this is not guaranteed to capture sampling irregularities during training. Alternatively, the Phased LSTM unit has been created to address this problem by updating its state using the sampling times explicitly. In this work, we study the effectiveness of the LSTM and Phased LSTM based architectures for the classification of astronomical lightcurves. We use seven catalogs containing periodic and nonperiodic astronomical objects. Our findings show that LSTM outperformed PLSTM on 6/7 datasets. However, the combination of both units enhances the results in all datasets.

astro-ph.IM

Alert Classification for the ALeRCE Broker System: The Light Curve Classifier

We present the first version of the ALeRCE (Automatic Learning for the Rapid Classification of Events) broker light curve classifier. ALeRCE is currently processing the Zwicky Transient Facility (ZTF) alert stream, in preparation for the Vera C. Rubin Observatory. The ALeRCE light curve classifier uses variability features computed from the ZTF alert stream, and colors obtained from AllWISE and ZTF photometry. We apply a Balanced Random Forest algorithm with a two-level scheme, where the top level classifies each source as periodic, stochastic, or transient, and the bottom level further resolves each of these hierarchical classes, amongst 15 total classes. This classifier corresponds to the first attempt to classify multiple classes of stochastic variables (including core- and host-dominated active galactic nuclei, blazars, young stellar objects, and cataclysmic variables) in addition to different classes of periodic and transient sources, using real data. We created a labeled set using various public catalogs (such as the Catalina Surveys and {\em Gaia} DR2 variable stars catalogs, and the Million Quasars catalog), and we classify all objects with $\geq6$ $g$-band or $\geq6$ $r$-band detections in ZTF (868,371 sources as of 2020/06/09), providing updated classifications for sources with new alerts every day. For the top level we obtain macro-averaged precision and recall scores of 0.96 and 0.99, respectively, and for the bottom level we obtain macro-averaged precision and recall scores of 0.57 and 0.76, respectively. Updated classifications from the light curve classifier can be found at the \href{http://alerce.online}{ALeRCE Explorer website}.

astro-ph.IM

A phylogenetic analysis of galaxies in the Coma Cluster and the field: a new approach to galaxy evolution

We propose a phylogenetic approach (PA) as a novel and robust tool to detect galaxy populations (GPs) based on their chemical composition. The branches of the tree are interpreted as different GPs and the length between nodes as the internal chemical variation along a branch. We apply the PA using 30 abundance indices from the Sloan Digital Sky Survey to 475 galaxies in the Coma Cluster and 438 galaxies in the field. We find that a dense environment, such as Coma, shows several GPs, which indicates that the environment is promoting galaxy evolution. Each population shares common properties that can be identified in colour magnitude space, in addition to minor structures inside the red sequence. The field is more homogeneous, presenting one main GP. We also apply a principal component analysis (PCA) to both samples, and find that the PCA does not have the same power in identifying GPs.

astro-ph.GA

The Automatic Learning for the Rapid Classification of Events (ALeRCE) Alert Broker

We introduce the Automatic Learning for the Rapid Classification of Events (ALeRCE) broker, an astronomical alert broker designed to provide a rapid and self--consistent classification of large etendue telescope alert streams, such as that provided by the Zwicky Transient Facility (ZTF) and, in the future, the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST). ALeRCE is a Chilean--led broker run by an interdisciplinary team of astronomers and engineers, working to become intermediaries between survey and follow--up facilities. ALeRCE uses a pipeline which includes the real--time ingestion, aggregation, cross--matching, machine learning (ML) classification, and visualization of the ZTF alert stream. We use two classifiers: a stamp--based classifier, designed for rapid classification, and a light--curve--based classifier, which uses the multi--band flux evolution to achieve a more refined classification. We describe in detail our pipeline, data products, tools and services, which are made public for the community (see \url{https://alerce.science}). Since we began operating our real--time ML classification of the ZTF alert stream in early 2019, we have grown a large community of active users around the globe. We describe our results to date, including the real--time processing of $9.7\times10^7$ alerts, the stamp classification of $1.9\times10^7$ objects, the light curve classification of $8.5\times10^5$ objects, the report of 3088 supernova candidates, and different experiments using LSST-like alert streams. Finally, we discuss the challenges ahead to go from a single-stream of alerts such as ZTF to a multi--stream ecosystem dominated by LSST.

astro-ph.IM

Asteroids' Size Distribution and Colors from HiTS

We report the observations of solar system objects during the 2015 campaign of the High cadence Transient Survey (HiTS). We found 5740 bodies (mostly Main Belt asteroids), 1203 of which were detected in different nights and in $g'$ and $r'$. Objects were linked in the barycenter system and their orbital parameters were computed assuming Keplerian motion. We identified 6 near Earth objects, 1738 Main Belt asteroids and 4 Trans-Neptunian objects. We did not find a $g'-r'$ color-size correlation for $14<H_{g'}<18$ ($1<D<10$ km) asteroids. We show asteroids' colors are disturbed by HiTS' 1.6 hour cadence and estimate that observations should be separated by at most 14 minutes to avoid confusion in future wide-field surveys like LSST. The size distribution for the Main Belt objects can be characterized as a simple power law with slope $\sim0.9$, steeper than in any other survey, while data from HiTS 2014's campaign is consistent with previous ones (slopes $\sim0.68$ at the bright end and $\sim0.34$ at the faint end). This difference is likely due to the ecliptic distribution of the Main Belt since 2015's campaign surveyed farther from the ecliptic than did 2014's and most previous surveys.

astro-ph.EP

The delay of shock breakout due to circumstellar material seen in most Type II Supernovae

Type II supernovae (SNe) originate from the explosion of hydrogen-rich supergiant massive stars. Their first electromagnetic signature is the shock breakout, a short-lived phenomenon which can last from hours to days depending on the density at shock emergence. We present 26 rising optical light curves of SN II candidates discovered shortly after explosion by the High cadence Transient Survey (HiTS) and derive physical parameters based on hydrodynamical models using a Bayesian approach. We observe a steep rise of a few days in 24 out of 26 SN II candidates, indicating the systematic detection of shock breakouts in a dense circumstellar matter consistent with a mass loss rate $\dot{M} > 10^{-4} M_\odot yr^{-1}$ or a dense atmosphere. This implies that the characteristic hour timescale signature of stellar envelope SBOs may be rare in nature and could be delayed into longer-lived circumstellar material shock breakouts in most Type II SNe.

astro-ph.HE

Asteroids in the High cadence Transient Survey

We report on the serendipitous observations of Solar System objects imaged during the High cadence Transient Survey (HiTS) 2014 observation campaign. Data from this high cadence, wide field survey was originally analyzed for finding variable static sources using Machine Learning to select the most-likely candidates. In this work we search for moving transients consistent with Solar System objects and derive their orbital parameters. We use a simple, custom detection algorithm to link trajectories and assume Keplerian motion to derive the asteroid's orbital parameters. We use known asteroids from the Minor Planet Center (MPC) database to assess the detection efficiency of the survey and our search algorithm. Trajectories have an average of nine detections spread over 2 days, and our fit yields typical errors of $σ_a\sim 0.07 ~{\rm AU}$, $σ_{\rm e} \sim 0.07 $ and $σ_i\sim 0.^{\circ}5~ {\rm deg}$ in semi-major axis, eccentricity, and inclination respectively for known asteroids in our sample. We extract 7,700 orbits from our trajectories, identifying 19 near Earth objects, 6,687 asteroids, 14 Centaurs, and 15 trans-Neptunian objects. This highlights the complementarity of supernova wide field surveys for Solar System research and the significance of machine learning to clean data of false detections. It is a good example of the data--driven science that LSST will deliver.

astro-ph.EP

A catalog of visual-like morphologies in the 5 CANDELS fields using deep-learning

We present a catalog of visual like H-band morphologies of $\sim50.000$ galaxies ($H_{f160w}<24.5$) in the 5 CANDELS fields (GOODS-N, GOODS-S, UDS, EGS and COSMOS). Morphologies are estimated with Convolutional Neural Networks (ConvNets). The median redshift of the sample is $ \sim1.25$. The algorithm is trained on GOODS-S for which visual classifications are publicly available and then applied to the other 4 fields. Following the CANDELS main morphology classification scheme, our model retrieves the probabilities for each galaxy of having a spheroid, a disk, presenting an irregularity, being compact or point source and being unclassifiable. ConvNets are able to predict the fractions of votes given a galaxy image with zero bias and $\sim10\%$ scatter. The fraction of miss-classifications is less than $1\%$. Our classification scheme represents a major improvement with respect to CAS (Concentration-Asymmetry-Smoothness)-based methods, which hit a $20-30\%$ contamination limit at high z. The catalog is released with the present paper via the $\href{http://rainbowx.fis.ucm.es/Rainbow_navigator_public}{Rainbow\,database}$

astro-ph.GA