SearcharxivSearch

arXiv subjects

Jeroen Audenaert

Publications and source records attributed to Jeroen Audenaert.

16 recordsLinked to original sources

Learning What's Real: Disentangling Signal and Measurement Artifacts in Multi-Sensor Data, with Applications to Astrophysics

Data collected from the physical world is always a combination of multiple sources: an underlying signal from the physical process of interest and a signal from measurement-dependent artifacts from the sensor or instrument. This secondary signal acts as a confounding factor, limiting our ability to extract information about the physics underlying the phenomena we observe. Furthermore, it complicates the combination of observations in heterogeneous or multi-instrument settings. We propose a deep learning framework that leverages overlapping observations, a dual-encoder architecture, and a counterfactual generation objective to disentangle these factors of variation. The resulting representations explicitly separate intrinsic signals from sensor-specific distortions and noise, and can be used for counterfactual view generation, parameter inference unconfounded by measurement distortions, and instrument-independent similarity search. We demonstrate the effectiveness of our approach on astrophysical galaxy images from the DESI Legacy Imaging Survey (Legacy) and the Hyper Suprime-Cam (HSC) Survey as a representative multi-instrument setting. This framework provides a general recipe for scientific and multi-modal self-supervised pretraining: construct training pairs from overlapping observations of the same physical system, treat sensor- or modality-specific effects as augmentations, and learn invariant representations through counterfactual generation.

astro-ph.IM

Variability classification of TESS targets in LOPS2, the first long-term pointing field of PLATO. Version 1 of the public variability catalogue

The PLAnetary Transits and Oscillations of stars (PLATO) mission is expected to launch in January 2027. A total of 8\% of its data rate will be dedicated to complementary science targets selected from approved Guest Observer proposals. We seek to provide an open-source catalogue of variable stars in PLATO's first long-term observing field, LOPS2. We want to use existing observations from the Transiting Exoplanet Survey Satellite (TESS), which has observed many stars in LOPS2. We classified 38 million calibrated aperture light curves from the TESS-Gaia Light Curve pipeline (TGLC, $G\lesssim17$) for 6 million unique sources in LOPS2 with two machine learning frameworks -- a deep neural network and a feature-based gradient-boosted decision-tree ensemble. We combined their predictions to create this first version of the LOPS2 variability catalogue, performed manual vetting of a sub-sample classified light curves, and a statistical analysis of the results to validate our methodology and to assess the variability properties and parameters of the stars in the catalogue. Our classification resulted in the identification of approximately 72% of the light curves having dominant instrument- or pipeline-induced signal, with the remaining 28% representing 3.6 million individual candidate variable stars, including pulsating, rotating, and eclipsing stars. Candidate pulsators exhibit varied behaviour in terms of their frequencies, amplitudes, rotation, and fundamental parameters. To ensure purity of the samples, filtering on colour, luminosity, the dominant frequency and its amplitude, and presence of close neighbours is helpful. We provide the first version of our PLATO LOPS2 variability catalogue to the community for further study and scrutiny. It is to date one of the largest catalogues of variable stars from an automated classification pipeline.

astro-ph.SR

ASTRAFier: A Novel and Scalable Transformer-based Stellar Variability Classifier

Photometric missions such as Kepler and TESS have generated millions of light curves covering almost the entire sky, offering unprecedented opportunities to study stellar variability and advance our understanding of the Universe. In this data-rich environment, machine learning has emerged as a powerful tool to efficiently and accurately process and classify light curves according to their type of stellar variability. In this work, we introduce ASTRAFier: a novel Transformer-based model for variability classification that integrates Bidirectional Long Short-Term Memory (BiLSTM) and Convolutional Neural Networks (CNNs). The model operates directly on time series without requiring feature engineering, creating an easy-to-maintain and efficient end-to-end classification framework. We train and validate our model using both Kepler and TESS light curves and, respectively, achieve a classification accuracy of $94.26\%$ on Kepler and $88.22\%$ on TESS. We demonstrate scalability by deploying our model on $\sim 2.8$ million TESS light curves from sectors 14, 15, and 26 (Kepler Field-of-View) delivered by MIT's Quick-look Pipeline (QLP) and release the resulting stellar variability catalog.

astro-ph.IM

The PLATO Science Calibration and Validation Plan: Targets for the First Long-pointing Field

In order to meet the science goals of the PLATO space mission, an extensive science calibration and validation plan has been designed. This paper describes this plan, as well as the methodology adopted to select the science calibration and validation stars that have entered its input catalogue. This is the so-called {\tt scvPIC}, which is part of the general PLATO Input Catalogue (PIC) for the first selected long pointing field in the Southern Hemisphere known as LOPS2. While many of PLATO's science requirements needed dedicated stars as calibrators as discussed here, its most stringent requirement is the delivery of the age of the host stars of exoplanetary systems with an accuracy better than 10\% for a G0V star of {\it V} = 10 mag, i.e. a nearby Sun-like star. This is presently not within reach for large populations of dwarfs and subgiants in the Milky Way as it requires the models of their stellar interiors to be improved. We discuss how this ambitious age requirement led to the selection of tens of thousands of red giants, and of thousands of main-sequence early F-type gravity-mode pulsators in order to deduce their internal rotation profile across stellar evolution. This asteroseismic observable will then be imported as key information into improved models of dwarfs and subgiants in the Milky Way as optimal modelling tools for ever better age-dating of the exoplanet hosts as the PLATO mission moves along. Additional calibrators and validators included in the {\tt scvPIC} are a few thousands of binaries, a few hundreds of legacy and benchmark stars, a few hundred photometrically stable stars, and six transiting brown dwarfs.

astro-ph.SR

QLP Data Release Notes 004: TESS-Gaia Light Curve Photometry Implementation

The Quick-Look Pipeline (QLP; Huang et al. 2020, Kunimoto et al. 2021 and references therein) generates light curves for up to 2 million stars every 27.4 days observed by TESS as part of its planet search. As machine learning methods enable deeper searches and scientific priorities shift toward fainter stars, there is a motivation for QLP to perform better at fainter magnitudes. We have adopted the photometry methods employed by the TESS-Gaia Light Curve package (Han & Brandt 2023), which has been shown to have better noise characteristics than the original QLP photometry from 10.5 $<$ $T$ $<$ 13.5. We still perform aperture photometry and deliver 3 apertures, and 3 levels of detrending for all stars brighter than $T$ = 13.5, so the changes should be seamless for external users. This method is implemented as of Sector 94 in QLP light curves and is providing users with higher precision light curves and allows detection of fainter signals in our planet searches.

astro-ph.EP

Astrometric follow-up of near-Earth asteroid 2024 YR4 during a Torino scale level 3 alert

The discovery of 2024 YR4 presented the planetary defense community with the most significant impact threat in almost two decades, reaching level 3 on the Torino scale. The community, now mature and well-organized, responded with a global observational effort. Astrometric measurements, forming the basis for orbital refinement and impact prediction, were a central component of this response. In this paper, we present the astrometric data collected by the international community, from the time of discovery until the object became too faint for all existing observational assets, including JWST. We also discuss the coordination role played by the International Asteroid Warning Network, and the importance of publicly available image archives to enable precovery searches.

astro-ph.EP

Simulation-Based Pretraining and Domain Adaptation for Astronomical Time Series with Minimal Labeled Data

Astronomical time-series analysis faces a critical limitation: the scarcity of labeled observational data. We present a pre-training approach that leverages simulations, significantly reducing the need for labeled examples from real observations. Our models, trained on simulated data from multiple astronomical surveys (ZTF and LSST), learn generalizable representations that transfer effectively to downstream tasks. Using classifier-based architectures enhanced with contrastive and adversarial objectives, we create domain-agnostic models that demonstrate substantial performance improvements over baseline methods in classification, redshift estimation, and anomaly detection when fine-tuned with minimal real data. Remarkably, our models exhibit effective zero-shot transfer capabilities, achieving comparable performance on future telescope (LSST) simulations when trained solely on existing telescope (ZTF) data. Furthermore, they generalize to very different astronomical phenomena (namely variable stars from NASA's \textit{Kepler} telescope) despite being trained on transient events, demonstrating cross-domain capabilities. Our approach provides a practical solution for building general models when labeled data is scarce, but domain knowledge can be encoded in simulations.

astro-ph.IM

Causal Foundation Models: Disentangling Physics from Instrument Properties

Foundation models for structured time series data must contend with a fundamental challenge: observations often conflate the true underlying physical phenomena with systematic distortions introduced by measurement instruments. This entanglement limits model generalization, especially in heterogeneous or multi-instrument settings. We present a causally-motivated foundation model that explicitly disentangles physical and instrumental factors using a dual-encoder architecture trained with structured contrastive learning. Leveraging naturally occurring observational triplets (i.e., where the same target is measured under varying conditions, and distinct targets are measured under shared conditions) our model learns separate latent representations for the underlying physical signal and instrument effects. Evaluated on simulated astronomical time series designed to resemble the complexity of variable stars observed by missions like NASA's Transiting Exoplanet Survey Satellite (TESS), our method significantly outperforms traditional single-latent space foundation models on downstream prediction tasks, particularly in low-data regimes. These results demonstrate that our model supports key capabilities of foundation models, including few-shot generalization and efficient adaptation, and highlight the importance of encoding causal structure into representation learning for structured data.

cs.LG

From stellar light to astrophysical insight: automating variable star research with machine learning

Large-scale photometric surveys are revolutionizing astronomy by delivering unprecedented amounts of data. The rich data sets from missions such as the NASA Kepler and TESS satellites, and the upcoming ESA PLATO mission, are a treasure trove for stellar variability, asteroseismology and exoplanet studies. In order to unlock the full scientific potential of these massive data sets, automated data-driven methods are needed. In this review, I illustrate how machine learning is bringing asteroseismology toward an era of automated scientific discovery, covering the full cycle from data cleaning to variability classification and parameter inference, while highlighting the recent advances in representation learning, multimodal datasets and foundation models. This invited review offers a guide to the challenges and opportunities machine learning brings for stellar variability research and how it could help unlock new frontiers in time-domain astronomy.

astro-ph.IM

A Disintegrating Rocky Planet with Prominent Comet-like Tails Around a Bright Star

We report the discovery of BD+05$\,$4868$\,$Ab, a transiting exoplanet orbiting a bright ($V=10.16$) K-dwarf (TIC 466376085) with a period of 1.27 days. Observations from NASA's Transiting Exoplanet Survey Satellite (TESS) reveal variable transit depths and asymmetric transit profiles that are characteristic of comet-like tails formed by dusty effluents emanating from a disintegrating planet. Unique to BD+05$\,$4868$\,$Ab is the presence of prominent dust tails in both the trailing and leading directions that contribute to the extinction of starlight from the host star. By fitting the observed transit profile and analytically modeling the drift of dust grains within both dust tails, we infer large grain sizes ($\sim1-10\,μ$m) and a mass loss rate of $10\,M_{\rm \oplus}\,$Gyr$^{-1}$, suggestive of a lunar-mass object with a disintegration timescale of only several Myr. The host star is probably older than the Sun and is accompanied by an M-dwarf companion at a projected physical separation of 130 AU. The brightness of the host star, combined with the planet's relatively deep transits ($0.8-2.0\%$), presents BD+05$\,$4868$\,$Ab as a prime target for compositional studies of rocky exoplanets and investigations into the nature of catastrophically evaporating planets.

astro-ph.EP

The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TB of Astronomical Scientific Data

We present the MULTIMODAL UNIVERSE, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, the MULTIMODAL UNIVERSE contains hundreds of millions of astronomical observations, constituting 100\,TB of multi-channel and hyper-spectral images, spectra, multivariate time series, as well as a wide variety of associated scientific measurements and "metadata". In addition, we include a range of benchmark tasks representative of standard practices for machine learning methods in astrophysics. This massive dataset will enable the development of large multi-modal models specifically targeted towards scientific applications. All codes used to compile the MULTIMODAL UNIVERSE and a description of how to access the data is available at https://github.com/MultimodalUniverse/MultimodalUniverse

astro-ph.IM

The PLATO Mission

PLATO (PLAnetary Transits and Oscillations of stars) is ESA's M3 mission designed to detect and characterise extrasolar planets and perform asteroseismic monitoring of a large number of stars. PLATO will detect small planets (down to <2 R_(Earth)) around bright stars (<11 mag), including terrestrial planets in the habitable zone of solar-like stars. With the complement of radial velocity observations from the ground, planets will be characterised for their radius, mass, and age with high accuracy (5 %, 10 %, 10 % for an Earth-Sun combination respectively). PLATO will provide us with a large-scale catalogue of well-characterised small planets up to intermediate orbital periods, relevant for a meaningful comparison to planet formation theories and to better understand planet evolution. It will make possible comparative exoplanetology to place our Solar System planets in a broader context. In parallel, PLATO will study (host) stars using asteroseismology, allowing us to determine the stellar properties with high accuracy, substantially enhancing our knowledge of stellar structure and evolution. The payload instrument consists of 26 cameras with 12cm aperture each. For at least four years, the mission will perform high-precision photometric measurements. Here we review the science objectives, present PLATO's target samples and fields, provide an overview of expected core science performance as well as a description of the instrument and the mission profile at the beginning of the serial production of the flight cameras. PLATO is scheduled for a launch date end 2026. This overview therefore provides a summary of the mission to the community in preparation of the upcoming operational phases.

astro-ph.IM

Applications of Deep Learning to physics workflows

Modern large-scale physics experiments create datasets with sizes and streaming rates that can exceed those from industry leaders such as Google Cloud and Netflix. Fully processing these datasets requires both sufficient compute power and efficient workflows. Recent advances in Machine Learning (ML) and Artificial Intelligence (AI) can either improve or replace existing domain-specific algorithms to increase workflow efficiency. Not only can these algorithms improve the physics performance of current algorithms, but they can often be executed more quickly, especially when run on coprocessors such as GPUs or FPGAs. In the winter of 2023, MIT hosted the Accelerating Physics with ML at MIT workshop, which brought together researchers from gravitational-wave physics, multi-messenger astrophysics, and particle physics to discuss and share current efforts to integrate ML tools into their workflows. The following white paper highlights examples of algorithms and computing frameworks discussed during this workshop and summarizes the expected computing needs for the immediate future of the involved fields.

hep-ex

Multiscale entropy analysis of astronomical time series. Discovering subclusters of hybrid pulsators

The multiscale entropy assesses the complexity of a signal across different timescales. It originates from the biomedical domain and was recently successfully used to characterize light curves as part of a supervised machine learning framework to classify stellar variability. We explore the behavior of the multiscale entropy in detail by studying its algorithmic properties in a stellar variability context and by linking it with traditional astronomical time series analysis methods. We subsequently use the multiscale entropy as the basis for an interpretable clustering framework that can distinguish hybrid pulsators with both p- and g-modes from stars with only p-mode pulsations, such as $δ$ Sct stars, or from stars with only g-mode pulsations, such as $γ$ Dor stars. We find that the multiscale entropy is a powerful tool for capturing variability patterns in stellar light curves. The multiscale entropy provides insights into the pulsation structure of a star and reveals how short- and long-term variability interact with each other based on time-domain information only. We also show that the multiscale entropy is correlated to the frequency content of a stellar signal and in particular to the near-core rotation rates of g-mode pulsators. We find that our new clustering framework can successfully identify the hybrid pulsators with both p- and g-modes in sets of $δ$ Sct and $γ$ Dor stars, respectively. The benefit of our clustering framework is that it is unsupervised. It therefore does not require previously labeled data and hence is not biased by previous knowledge.

astro-ph.SR

Zeta-Payne: a fully automated spectrum analysis algorithm for the Milky Way Mapper program of the SDSS-V survey

The Sloan Digital Sky Survey has recently initiated its 5th survey generation (SDSS-V), with a central focus on stellar spectroscopy. In particular, SDSS-V Milky Way Mapper program will deliver multi-epoch optical and near-infrared spectra for more than 5 million stars across the entire sky, covering a large range in stellar mass, surface temperature, evolutionary stage, and age. About 10% of those spectra will be of hot stars of OBAF spectral types, for whose analysis no established survey pipelines exist. Here we present the spectral analysis algorithm, Zeta-Payne, developed specifically to obtain stellar labels from SDSS-V spectra of stars with these spectral types and drawing on machine learning tools. We provide details of the algorithm training, its test on artificial spectra, and its validation on two control samples of real stars. Analysis with Zeta-Payne leads to only modest internal uncertainties in the near-IR with APOGEE (optical with BOSS): 3-10% (1-2%) for Teff, 5-30% (5-25%) for v*sin(i), 1.7-6.3 km/s(0.7-2.2 km/s) for RV, $<0.1$ dex ($<0.05$ dex) for log(g), and 0.4-0.5 dex (0.1 dex) for [M/H] of the star, respectively. We find a good agreement between atmospheric parameters of OBAF-type stars when inferred from their high- and low-resolution optical spectra. For most stellar labels the APOGEE spectra are (far) less informative than the BOSS spectra of these stars, while log(g), v*sin(i), and [M/H] are in most cases too uncertain for meaningful astrophysical interpretation. This makes BOSS low-resolution optical spectra better for stellar labels of OBAF-type stars, unless the latter are subject to high levels of extinction.

astro-ph.IM

TESS Data for Asteroseismology (T'DA) Stellar Variability Classification Pipeline: Set-Up and Application to the Kepler Q9 Data

The NASA Transiting Exoplanet Survey Satellite (TESS) is observing tens of millions of stars with time spans ranging from $\sim$ 27 days to about 1 year of continuous observations. This vast amount of data contains a wealth of information for variability, exoplanet, and stellar astrophysics studies but requires a number of processing steps before it can be fully utilized. In order to efficiently process all the TESS data and make it available to the wider scientific community, the TESS Data for Asteroseismology working group, as part of the TESS Asteroseismic Science Consortium, has created an automated open-source processing pipeline to produce light curves corrected for systematics from the short- and long-cadence raw photometry data and to classify these according to stellar variability type. We will process all stars down to a TESS magnitude of 15. This paper is the next in a series detailing how the pipeline works. Here, we present our methodology for the automatic variability classification of TESS photometry using an ensemble of supervised learners that are combined into a metaclassifier. We successfully validate our method using a carefully constructed labelled sample of Kepler Q9 light curves with a 27.4 days time span mimicking single-sector TESS observations, on which we obtain an overall accuracy of 94.9%. We demonstrate that our methodology can successfully classify stars outside of our labeled sample by applying it to all $\sim$ 167,000 stars observed in Q9 of the Kepler space mission.

astro-ph.SR