SearcharxivSearch

arXiv subjects

Pablo Huijse

Publications and source records attributed to Pablo Huijse.

At least 19 recordsLinked to original sources

Plato's view on supermassive black hole binaries: Exploring the faint limit of ESA's Plato space mission

The search for supermassive black hole binaries (SMBHBs) has, in recent years, seen the dawn of exploration with several hundred candidates claimed from photometric and spectroscopic surveys monitoring AGNs. While only a handful persist to date, the advent of upcoming high-precision wide-field photometric missions motivates continuing the pursuit of confirming SMBHBs in the optical. We explore the possibility of using the ESA Plato space mission to detect photometric signatures of SMBHBs. Motivated by the Kepler observation of Spikey, the best known self-lensing flare (SLF) candidate to date, this work aims to benchmark the scientific outcome if Plato were to observe Spikey-like objects via its Guest Observer programme. Starting from the Gaia database, we assemble a catalogue of 12,226 bright ($G < 19$) high-probability Quasars for the two pointing fields of Plato's nominal mission. This Plato Quasar catalogue will be pivotal for future follow-up observations of larger photometric searches such as the Vera Rubin LSST survey. We use the Plato camera simulator, PlatoSim, to realistically explore the noise budget in Plato's faint limit, while generating mock light curves to benchmark Plato's ability to recover signatures of SMBHBs. We show that, although not at all designed for the purpose, Plato is capable of detecting Spikey-like SMBHB candidates through their relativistic photometric signatures using Bayesian inference and evidence. Plato will in particular be able to confirm or rule out Spikey and Spikey-like objects with a limiting magnitude of $G\leq18$. With a minimum 2-yr baseline per pointing field, we show that Plato not only could play an essential role in future SMBHB research, but may be an integrated part of the observational fleet of continuous high-precision facilities monitoring SMBHB candidates in the near future.

astro-ph.GA

Variability classification of TESS targets in LOPS2, the first long-term pointing field of PLATO. Version 1 of the public variability catalogue

The PLAnetary Transits and Oscillations of stars (PLATO) mission is expected to launch in January 2027. A total of 8\% of its data rate will be dedicated to complementary science targets selected from approved Guest Observer proposals. We seek to provide an open-source catalogue of variable stars in PLATO's first long-term observing field, LOPS2. We want to use existing observations from the Transiting Exoplanet Survey Satellite (TESS), which has observed many stars in LOPS2. We classified 38 million calibrated aperture light curves from the TESS-Gaia Light Curve pipeline (TGLC, $G\lesssim17$) for 6 million unique sources in LOPS2 with two machine learning frameworks -- a deep neural network and a feature-based gradient-boosted decision-tree ensemble. We combined their predictions to create this first version of the LOPS2 variability catalogue, performed manual vetting of a sub-sample classified light curves, and a statistical analysis of the results to validate our methodology and to assess the variability properties and parameters of the stars in the catalogue. Our classification resulted in the identification of approximately 72% of the light curves having dominant instrument- or pipeline-induced signal, with the remaining 28% representing 3.6 million individual candidate variable stars, including pulsating, rotating, and eclipsing stars. Candidate pulsators exhibit varied behaviour in terms of their frequencies, amplitudes, rotation, and fundamental parameters. To ensure purity of the samples, filtering on colour, luminosity, the dominant frequency and its amplitude, and presence of close neighbours is helpful. We provide the first version of our PLATO LOPS2 variability catalogue to the community for further study and scrutiny. It is to date one of the largest catalogues of variable stars from an automated classification pipeline.

astro-ph.SR

Automated all-sky detection of {\gamma} Doradus / {\delta} Scuti hybrids in TESS data from positive unlabelled (PU) learning

The Transiting Exoplanet Survey Satellite (TESS) mission has observed hundreds of millions of stars, substantially contributing to the available pool of high-precision photometric space data. Among them are the relatively rare $\gamma$ Doradus / $\delta$ Scuti ($\gamma$ Dor / $\delta$ Sct) hybrid pulsators, which have been previously studied using Kepler data. These stars are perfect laboratories to probe both inner and outer interior stellar layers thanks to them exhibiting both pressure and gravity modes. We seek to classify an all-sky sample of AF stars observed by TESS to find previously undiscovered hybrid pulsators and supply them in a catalogue of candidates. We also aim to compare the light curves produced with the TESS-Gaia Light Curve (TGLC) pipeline, currently underused in variability studies, with other publicly available light curves. We compared dominant and secondary frequencies of confirmed hybrid pulsators in Kepler, extended mission Quick Look Pipeline (QLP) data, and nominal and extended mission TGLC data. We then used a feature-based positive unlabelled (PU) learning classifier to search for new hybrid pulsators amongst TESS AF stars and investigated the properties of the detected populations. We find that the variability of confirmed hybrids in TGLC agrees well with the one occurring in QLP light curves and has a high recovery rate of \kepler-extracted frequencies. Our `smart binning' method allows for robust extraction of hybrids from large unlabelled datasets, with an average out-of-bag prediction for test set hybrids at 93.04\%. The analysis of dominant frequencies in high-probability candidates shows that we find more pressure-mode dominant hybrids. Our catalogue includes 62,026 new candidate light curves from the nominal and extended TESS missions, with individual probabilities of being a hybrid in each available sector.

astro-ph.SR

A search for periodic AGN variability in $\textit{Gaia}$ Data Release 3

Supermassive black hole binaries (SMBHB) are expected to produce periodic modulations in active galactic nuclei (AGN) light curves, but distinguishing such signals from stochastic red-noise variability remains a major challenge. We present the first systematic search for statistically significant AGN periodicities using the optical photometry from the Gaia space mission Data Release 3 (DR3), with the goal of identifying SMBHB candidates and establishing a methodological data analysis framework that can be scaled to the forthcoming Data Release 4 (DR4). We analyse Gaia G band light curves of 377,128 sources from the Gaia celestial reference frame (CRF3). Stochastic variability is modelled as a damped random walk Gaussian process, and empirical false alarm probabilities are derived by comparing observed Lomb-Scargle periodogram peaks against 100,000 synthetic red-noise realisations. Candidates from this first stage are then re-evaluated using full Markov chain Monte Carlo inference under both exponential and powered-exponential kernels. We find 13 sources surviving our statistical criterion ($p < \alpha = 10^{-5}$) after both stages of filtering, which is consistent with the expected false-positive rate. All candidates cover fewer than 2.5 cycles of the candidate period and are systematically concentrated in a region of the parameter space indicative of model misspecification. No reliable periodic SMBHB candidates are retained. The ${\sim}950$-day baseline of Gaia DR3 confines all detections to the few-cycle regime where red noise most convincingly mimics periodicity, a limitation that photometric precision alone cannot overcome. The longer baseline of Gaia DR4 will be essential to push beyond this regime. We offer our data analysis software pipeline in open access to the community.

astro-ph.HE

A self-regulated convolutional neural network for classifying variable stars

Over the last two decades, machine learning models have been widely applied and have proven effective in classifying variable stars, particularly with the adoption of deep learning architectures such as convolutional neural networks, recurrent neural networks, and transformer models. While these models have achieved high accuracy, they require high-quality, representative data and a large number of labelled samples for each star type to generalise well, which can be challenging in time-domain surveys. This challenge often leads to models learning and reinforcing biases inherent in the training data, an issue that is not easily detectable when validation is performed on subsamples from the same catalogue. The problem of biases in variable star data has been largely overlooked, and a definitive solution has yet to be established. In this paper, we propose a new approach to improve the reliability of classifiers in variable star classification by introducing a self-regulated training process. This process utilises synthetic samples generated by a physics-enhanced latent space variational autoencoder, incorporating six physical parameters from Gaia Data Release 3. Our method features a dynamic interaction between a classifier and a generative model, where the generative model produces ad-hoc synthetic light curves to reduce confusion during classifier training and populate underrepresented regions in the physical parameter space. Experiments conducted under various scenarios demonstrate that our self-regulated training approach outperforms traditional training methods for classifying variable stars on biased datasets, showing statistically significant improvements.

cs.LG

Informative regularization for a multi-layer perceptron RR Lyrae classifier under data shift

In recent decades, machine learning has provided valuable models and algorithms for processing and extracting knowledge from time-series surveys. Different classifiers have been proposed and performed to an excellent standard. Nevertheless, few papers have tackled the data shift problem in labeled training sets, which occurs when there is a mismatch between the data distribution in the training set and the testing set. This drawback can damage the prediction performance in unseen data. Consequently, we propose a scalable and easily adaptable approach based on an informative regularization and an ad-hoc training procedure to mitigate the shift problem during the training of a multi-layer perceptron for RR Lyrae classification. We collect ranges for characteristic features to construct a symbolic representation of prior knowledge, which was used to model the informative regularizer component. Simultaneously, we design a two-step back-propagation algorithm to integrate this knowledge into the neural network, whereby one step is applied in each epoch to minimize classification error, while another is applied to ensure regularization. Our algorithm defines a subset of parameters (a mask) for each loss function. This approach handles the forgetting effect, which stems from a trade-off between these loss functions (learning from data versus learning expert knowledge) during training. Experiments were conducted using recently proposed shifted benchmark sets for RR Lyrae stars, outperforming baseline models by up to 3\% through a more reliable classifier. Our method provides a new path to incorporate knowledge from characteristic features into artificial neural networks to manage the underlying data shift problem.

astro-ph.IM

DELIGHT: Deep Learning Identification of Galaxy Hosts of Transients using Multi-resolution Images

We present DELIGHT, or Deep Learning Identification of Galaxy Hosts of Transients, a new algorithm designed to automatically and in real-time identify the host galaxies of extragalactic transients. The proposed algorithm receives as input compact, multi-resolution images centered at the position of a transient candidate and outputs two-dimensional offset vectors that connect the transient with the center of its predicted host. The multi-resolution input consists of a set of images with the same number of pixels, but with progressively larger pixel sizes and fields of view. A sample of \nSample galaxies visually identified by the ALeRCE broker team was used to train a convolutional neural network regression model. We show that this method is able to correctly identify both relatively large ($10\arcsec < r < 60\arcsec$) and small ($r \le 10\arcsec$) apparent size host galaxies using much less information (32 kB) than with a large, single-resolution image (920 kB). The proposed method has fewer catastrophic errors in recovering the position and is more complete and has less contamination ($< 0.86\%$) recovering the cross-matched redshift than other state-of-the-art methods. The more efficient representation provided by multi-resolution input images could allow for the identification of transient host galaxies in real-time, if adopted in alert streams from new generation of large etendue telescopes such as the Vera C. Rubin Observatory.

astro-ph.IM

Bayesian Reconstruction of Fourier Pairs

In a number of data-driven applications such as detection of arrhythmia, interferometry or audio compression, observations are acquired indistinctly in the time or frequency domains: temporal observations allow us to study the spectral content of signals (e.g., audio), while frequency-domain observations are used to reconstruct temporal/spatial data (e.g., MRI). Classical approaches for spectral analysis rely either on i) a discretisation of the time and frequency domains, where the fast Fourier transform stands out as the \textit{de facto} off-the-shelf resource, or ii) stringent parametric models with closed-form spectra. However, the general literature fails to cater for missing observations and noise-corrupted data. Our aim is to address the lack of a principled treatment of data acquired indistinctly in the temporal and frequency domains in a way that is robust to missing or noisy observations, and that at the same time models uncertainty effectively. To achieve this aim, we first define a joint probabilistic model for the temporal and spectral representations of signals, to then perform a Bayesian model update in the light of observations, thus jointly reconstructing the complete (latent) time and frequency representations. The proposed model is analysed from a classical spectral analysis perspective, and its implementation is illustrated through intuitive examples. Lastly, we show that the proposed model is able to perform joint time and frequency reconstruction of real-world audio, healthcare and astronomy signals, while successfully dealing with missing data and handling uncertainty (noise) naturally against both classical and modern approaches for spectral estimation.

eess.SP

MPCC: Matching Priors and Conditionals for Clustering

Clustering is a fundamental task in unsupervised learning that depends heavily on the data representation that is used. Deep generative models have appeared as a promising tool to learn informative low-dimensional data representations. We propose Matching Priors and Conditionals for Clustering (MPCC), a GAN-based model with an encoder to infer latent variables and cluster categories from data, and a flexible decoder to generate samples from a conditional latent space. With MPCC we demonstrate that a deep generative model can be competitive/superior against discriminative methods in clustering tasks surpassing the state of the art over a diverse set of benchmark datasets. Our experiments show that adding a learnable prior and augmenting the number of encoder updates improve the quality of the generated samples, obtaining an inception score of 9.49 $\pm$ 0.15 and improving the Fréchet inception distance over the state of the art by a 46.9% in CIFAR10.

cs.LG

An Information Theory Approach on Deciding Spectroscopic Follow Ups

Classification and characterization of variable phenomena and transient phenomena are critical for astrophysics and cosmology. These objects are commonly studied using photometric time series or spectroscopic data. Given that many ongoing and future surveys are in time-domain and given that adding spectra provide further insights but requires more observational resources, it would be valuable to know which objects should we prioritize to have spectrum in addition to time series. We propose a methodology in a probabilistic setting that determines a-priory which objects are worth taking spectrum to obtain better insights, where we focus 'insight' as the type of the object (classification). Objects for which we query its spectrum are reclassified using their full spectrum information. We first train two classifiers, one that uses photometric data and another that uses photometric and spectroscopic data together. Then for each photometric object we estimate the probability of each possible spectrum outcome. We combine these models in various probabilistic frameworks (strategies) which are used to guide the selection of follow up observations. The best strategy depends on the intended use, whether it is getting more confidence or accuracy. For a given number of candidate objects (127, equal to 5% of the dataset) for taking spectra, we improve 37% class prediction accuracy as opposed to 20% of a non-naive (non-random) best base-line strategy. Our approach provides a general framework for follow-up strategies and can be extended beyond classification and to include other forms of follow-ups beyond spectroscopy.

astro-ph.IM

Deep Learning for Image Sequence Classification of Astronomical Events

We propose a new sequential classification model for astronomical objects based on a recurrent convolutional neural network (RCNN) which uses sequences of images as inputs. This approach avoids the computation of light curves or difference images. This is the first time that sequences of images are used directly for the classification of variable objects in astronomy. The second contribution of this work is the image simulation process. We generate synthetic image sequences that take into account the instrumental and observing conditions, obtaining a realistic, set of movies for each astronomical object. The simulated dataset is used to train our RCNN classifier. This approach allows us to generate datasets to train and test our RCNN model for different astronomical surveys and telescopes. We aim at building a simulated dataset whose distribution is close enough to the real dataset, so that a fine tuning could match the distributions between real and simulated dataset. To test the RCNN classifier trained with the synthetic dataset, we used real-world data from the High cadence Transient Survey (HiTS) obtaining an average recall of 85%, improved to 94% after performing fine tuning with 10 real samples per class. We compare the results of our model with those of a light curve random forest classifier. The proposed RCNN with fine tuning has a similar performance on the HiTS dataset compared to the light curve classifier, trained on an augmented training set with 10 real samples per class. The RCNN approach presents several advantages in an alert stream classification scenario, such as a reduction of the data pre-processing, faster online evaluation and easier performance improvement using a few real data samples. These results encourage us to use this method for alert brokers systems that will process alert streams generated by new telescopes such as the Large Synoptic Survey Telescope.

astro-ph.IM

The High Cadence Transient Survey (HITS): Compilation and characterization of light-curve catalogs

The High Cadence Transient Survey (HiTS) aims to discover and study transient objects with characteristic timescales between hours and days, such as pulsating, eclipsing and exploding stars. This survey represents a unique laboratory to explore large etendue observations from cadences of about 0.1 days and to test new computational tools for the analysis of large data. This work follows a fully \textit{Data Science} approach: from the raw data to the analysis and classification of variable sources. We compile a catalog of ${\sim}15$ million object detections and a catalog of ${\sim}2.5$ million light-curves classified by variability. The typical depth of the survey is $24.2$, $24.3$, $24.1$ and $23.8$ in $u$, $g$, $r$ and $i$ bands, respectively. We classified all point-like non-moving sources by first extracting features from their light-curves and then applying a Random Forest classifier. For the classification, we used a training set constructed using a combination of cross-matched catalogs, visual inspection, transfer/active learning and data augmentation. The classification model consists of several Random Forest classifiers organized in a hierarchical scheme. The classifier accuracy estimated on a test set is approximately $97\%$. In the unlabeled data, $3\,485$ sources were classified as variables, of which $1\,321$ were classified as periodic. Among the periodic classes we discovered with high confidence, 1 $δ$-scutti, 39 eclipsing binaries, 48 rotational variables and 90 RR-Lyrae and for the non-periodic classes we discovered 1 cataclysmic variables, 630 QSO, and 1 supernova candidates. The first data release can be accessed in the project archive of HiTS.

astro-ph.IM

Enhanced Rotational Invariant Convolutional Neural Network for Supernovae Detection

In this paper, we propose an enhanced CNN model for detecting supernovae (SNe). This is done by applying a new method for obtaining rotational invariance that exploits cyclic symmetry. In addition, we use a visualization approach, the layer-wise relevance propagation (LRP) method, which allows finding the relevant pixels in each image that contribute to discriminate between SN candidates and artifacts. We introduce a measure to assess quantitatively the effect of the rotational invariant methods on the LRP relevance heatmaps. This allows comparing the proposed method, CAP, with the original Deep-HiTS model. The results show that the enhanced method presents an augmented capacity for achieving rotational invariance with respect to the original model. An ensemble of CAP models obtained the best results so far on the HiTS dataset, reaching an average accuracy of 99.53%. The improvement over Deep-HiTS is significant both statistically and in practice.

astro-ph.IM

The VVV Survey RR Lyrae Population in the Galactic Centre Region

Deep near-IR images from the VVV Survey were used to search for RR Lyrae type ab (RRab) stars within 100' from the Galactic Centre (GC). A sample of 960 RRab stars were discovered. We use the reddening-corrected magnitudes in order to isolate RRab belonging to the GC. The mean period for our RRab sample is $P=0.5446$ days, yielding a mean metallicity of $[Fe/H] = -1.30$ dex and a median distance from the Sun of $D=8.05$. We measure the RRab surface density using the less reddened region sampled here, finding $1000$ RRab/sq deg at a projected Galactocentric distance $R_G=1.6$ deg. This implies a large total mass ($M>10^9 M_\odot$) for the old and metal-poor population contained inside $R_G$. We measure accurate relative proper motions, from which we derive tangential velocity dispersions of $σV_l = 125.0$ and $σV_b = 124.1$ km/s along the Galactic longitude and latitude coordinates, respectively. The fact that these quantities are similar indicate that the bulk rotation of the RRab population is negligible, and implies that this population is supported by velocity dispersion. There are two main conclusions of this study. First, the population as a whole is no different from the outer bulge RRab, predominantly a metal-poor component that is shifted respect the Oosterhoff type I population defined by the globular clusters in the halo. Second, the RRab sample, as representative of the old and metal-poor stellar population in the region, have high velocity dispersions and zero rotation, suggesting a formation via dissipational collapse.

astro-ph.GA

Robust period estimation using mutual information for multi-band light curves in the synoptic survey era

The Large Synoptic Survey Telescope (LSST) will produce an unprecedented amount of light curves using six optical bands. Robust and efficient methods that can aggregate data from multidimensional sparsely-sampled time series are needed. In this paper we present a new method for light curve period estimation based on the quadratic mutual information (QMI). The proposed method does not assume a particular model for the light curve nor its underlying probability density and it is robust to non-Gaussian noise and outliers. By combining the QMI from several bands the true period can be estimated even when no single-band QMI yields the period. Period recovery performance as a function of average magnitude and sample size is measured using 30,000 synthetic multi-band light curves of RR Lyrae and Cepheid variables generated by the LSST Operations and Catalog simulators. The results show that aggregating information from several bands is highly beneficial in LSST sparsely-sampled time series, obtaining an absolute increase in period recovery rate up to 50%. We also show that the QMI is more robust to noise and light curve length (sample size) than the multiband generalizations of the Lomb Scargle and Analysis of Variance periodograms, recovering the true period in 10-30% more cases than its competitors. A python package containing efficient Cython implementations of the QMI and other methods is provided.

astro-ph.IM

The High Cadence Transient Survey (HiTS) - I. Survey design and supernova shock breakout constraints

We present the first results of the High cadence Transient Survey (HiTS), a survey whose objective is to detect and follow up optical transients with characteristic timescales from hours to days, especially the earliest hours of supernova (SN) explosions. HiTS uses the Dark Energy Camera (DECam) and a custom made pipeline for image subtraction, candidate filtering and candidate visualization, which runs in real-time to be able to react rapidly to the new transients. We discuss the survey design, the technical challenges associated with the real-time analysis of these large volumes of data and our first results. In our 2013, 2014 and 2015 campaigns we have detected more than 120 young SN candidates, but we did not find a clear signature from the short-lived SN shock breakouts (SBOs) originating after the core collapse of red supergiant stars, which was the initial science aim of this survey. Using the empirical distribution of limiting-magnitudes from our observational campaigns we measured the expected recovery fraction of randomly injected SN light curves which included SBO optical peaks produced with models from Tominaga et al. (2011) and Nakar & Sari (2010). From this analysis we cannot rule out the models from Tominaga et al. (2011) under any reasonable distributions of progenitor masses, but we can marginally rule out the brighter and longer-lived SBO models from Nakar & Sari (2010) under our best-guess distribution of progenitor masses. Finally, we highlight the implications of this work for future massive datasets produced by astronomical observatories such as LSST.

astro-ph.SR

Computational Intelligence Challenges and Applications on Large-Scale Astronomical Time Series Databases

Time-domain astronomy (TDA) is facing a paradigm shift caused by the exponential growth of the sample size, data complexity and data generation rates of new astronomical sky surveys. For example, the Large Synoptic Survey Telescope (LSST), which will begin operations in northern Chile in 2022, will generate a nearly 150 Petabyte imaging dataset of the southern hemisphere sky. The LSST will stream data at rates of 2 Terabytes per hour, effectively capturing an unprecedented movie of the sky. The LSST is expected not only to improve our understanding of time-varying astrophysical objects, but also to reveal a plethora of yet unknown faint and fast-varying phenomena. To cope with a change of paradigm to data-driven astronomy, the fields of astroinformatics and astrostatistics have been created recently. The new data-oriented paradigms for astronomy combine statistics, data mining, knowledge discovery, machine learning and computational intelligence, in order to provide the automated and robust methods needed for the rapid detection and classification of known astrophysical objects as well as the unsupervised characterization of novel phenomena. In this article we present an overview of machine learning and computational intelligence applications to TDA. Future big data challenges and new lines of research in TDA, focusing on the LSST, are identified and discussed from the viewpoint of computational intelligence/machine learning. Interdisciplinary collaboration will be required to cope with the challenges posed by the deluge of astronomical data coming from the LSST.

astro-ph.IM

A Novel, Fully Automated Pipeline for Period Estimation in the EROS 2 Data Set

We present a new method to discriminate periodic from non-periodic irregularly sampled lightcurves. We introduce a periodic kernel and maximize a similarity measure derived from information theory to estimate the periods and a discriminator factor. We tested the method on a dataset containing 100,000 synthetic periodic and non-periodic lightcurves with various periods, amplitudes and shapes generated using a multivariate generative model. We correctly identified periodic and non-periodic lightcurves with a completeness of 90% and a precision of 95%, for lightcurves with a signal-to-noise ratio (SNR) larger than 0.5. We characterize the efficiency and reliability of the model using these synthetic lightcurves and applied the method on the EROS-2 dataset. A crucial consideration is the speed at which the method can be executed. Using hierarchical search and some simplification on the parameter search we were able to analyze 32.8 million lightcurves in 18 hours on a cluster of GPGPUs. Using the sensitivity analysis on the synthetic dataset, we infer that 0.42% in the LMC and 0.61% in the SMC of the sources show periodic behavior. The training set, the catalogs and source code are all available in http://timemachine.iic.harvard.edu.

astro-ph.IM