Searcharxiv⌕ Search

arXiv subjects

Nikolaos Nikolaou

Publications and source records attributed to Nikolaos Nikolaou.

At least 19 recordsLinked to original sources

aRieL - Reinforcement Learning for the Ariel Space Telescope Scheduling

The Ariel mission will conduct a population level survey of hundreds of exoplanet atmospheres, requiring observations to be scheduled across a large and diverse catalogue of potential targets. By framing Ariel scheduling as a long-horizon optimisation problem in which individual decisions affect both mission time and the scientific opportunities available later in the survey, Reinforcement Learning (RL) can be utilised as a framework for target selection. We introduce aRieL, a simulated observing environment coupled to a Set Transformer Proximal Policy Optimisation (PPO) agent that is trained to generate a scheduling policy. The RL policy consistently finds a strong compromise between competing survey objectives, producing large Tier~1 samples while maintaining high Tier~3 completion, population coverage, and observing efficiency when compared to a set of baseline heuristic polices. This behaviour is retained when the candidate catalogue is substantially expanded, while modifications to the reward function produce corresponding changes in the learned observing strategy. We further show that a policy trained on the baseline mission scenario can generalise without retraining to a modified scenario requiring repeated Tier~3 observations, revealing the resulting trade-off with the wider survey. These results demonstrate that reinforcement learning provides a flexible approach to large-scale astronomical scheduling, in which the observing strategy can respond directly to changes in mission state and to the scientific priorities, and motivate its wider exploration for observatories with similarly complex, state dependent scheduling problems.

astro-ph.EP↗

Graph neural networks for exoplanet atmospheres

Calculating disequilibrium chemistry in exoplanet atmospheres remains a significant computational bottleneck in atmospheric retrievals. The increasing observational precision from facilities such as JWST and the Ariel mission requires including disequilibrium chemistry in these analyses. Previous studies have demonstrated that neural networks can emulate kinetic chemistry, although their spatial inductive bias does not align with the topology of chemical reaction networks. This study introduces a graph neural network surrogate that represents chemical species as nodes and temperature-dependent reaction rates as edges, thereby enabling information propagation along physically meaningful chemical pathways. The model is trained on atmospheres generated using the Venot+2020 chemical scheme and Guillot temperature-pressure profiles. The GNN accurately reconstructs disequilibrium abundances across the sampled parameter space and reduces the mean abundance error by a factor of approximately 3 compared to the previous U-Net model. When applied to transmission spectra, most predictions fall within the observational precision expected for JWST and Ariel, with only about 7% of test atmospheres exceeding a 20 ppm mean spectral error. Performance variations are primarily observed in chemically transitional regimes near a carbon-to-oxygen ratio of one and at low temperatures. An evaluation of the boundary-case planet WASP-39b demonstrates effective performance under a moderate domain shift. Perturbation analysis indicates that disturbances propagate along chemical connectivity rather than spatial adjacency, confirming that the architecture captures the structure of reaction networks. These results suggest that GNN surrogates provide accurate, computationally efficient predictions of disequilibrium chemistry, facilitating integration into the atmospheric retrieval pipeline.

astro-ph.IM↗

PERTURB-c: Correlation Aware Perturbation Explainability for Regression Techniques to Understand Retrieval Black-boxes

In this paper we introduce PERTURB-c, a correlation-aware framework for interpreting black box regression models with one-dimensional structured inputs. We demonstrate this framework on a simulated case study with machine learning based transit spectroscopy retrievals of exoplanet WASP-107b. Characterising many exoplanet atmospheres can answer important questions about planetary populations, but traditional retrievals are very resource intensive; machine learning based methods offer a fast alternative however (i) they require high volumes data (only obtainable through simulations) to train and (ii) their complexity renders them black-boxes. Better understanding how they reach predictions can allow us to inspect for biases, which is especially important with simulated data, and verify that predictions are made on the basis of physically plausible features. This ultimately improves the ease of adoption of machine learning techniques. The most used methods to explain machine learning model predictions (such as SHAP and other methods that rely on stochastic sample generation) suffer from high computational complexity and struggle to account for interactions between inputs. PERTURB-c addresses these issues by leveraging physical knowledge of the known spectral correlation. For visualisation of this analysis, we propose a heat-map-based representation which is better suited to large numbers of input features along a single dimension, and that is more intuitive to those who are already familiar with retrieval methods. Note that while we chose this exoplanet retrieval context to demonstrate our methodologies, the PERTURB-c framework is model agnostic and in a broader context has potential value across a plethora of adjacent regression problems.

astro-ph.EP↗

Adaptive Online Emulation for Accelerating Complex Physical Simulations

Complex physical simulations often require trade-offs between model fidelity and computational feasibility. We introduce Adaptive Online Emulation (AOE), which dynamically learns neural network surrogates during simulation execution to accelerate expensive components. Unlike existing methods requiring extensive offline training, AOE uses Online Sequential Extreme Learning Machines (OS-ELMs) to continuously adapt emulators along the actual simulation trajectory. We employ a numerically stable variant of the OS-ELM using cumulative sufficient statistics to avoid matrix inversion instabilities. AOE integrates with time-stepping frameworks through a three-phase strategy balancing data collection, updates, and surrogate usage, while requiring orders of magnitude less training data than conventional surrogate approaches. Demonstrated on a 1D atmospheric model of exoplanet GJ1214b, AOE achieves 11.1 times speedup (91% time reduction) across 200,000 timesteps while maintaining accuracy, potentially making previously intractable high-fidelity time-stepping simulations computationally feasible.

physics.comp-ph↗

Planetary Edge Trends (PET). I. The Inner Edge-Stellar Mass Correlation

The position of the innermost planet (i.e., the inner edge) in a planetary system provides important information about the relationship of the entire system to its host star properties, offering potentially valuable insights into planetary formation and evolution processes. In this work, based on the Kepler Data Release 25 (DR25) catalog combined with LAMOST and Gaia data, we investigate the correlation between stellar mass and the inner edge position across different populations of small planets in multi-planetary systems, such as super-Earths and sub-Neptunes. By correcting for the influence of stellar metallicity and analyzing the impact of observational selection effects, we confirm the trend that as stellar mass increases, the position of the inner edge shifts outward. Our results reveal a stronger correlation between the inner edge and stellar mass with a power-law index of 0.6-1.1, which is larger compared to previous studies. The stronger correlation in our findings is primarily attributed to two factors: first, the metallicity correction applied in this work enhances the correlation; second, the previous use of occurrence rates to trace the inner edge weakens the observed correlation. Through comparison between observed statistical results and current theoretical models, we find that the pre-main-sequence (PMS) dust sublimation radius of the protoplanetary disk best matches the observed inner edge stellar mass. Therefore, we conclude that the inner dust disk likely limits the innermost orbits of small planets, contrasting with the inner edges of hot Jupiters, which are associated with the magnetospheres of gas disks, as suggested by previous studies. This highlights that the inner edges of different planetary populations are likely regulated by distinct mechanisms.

astro-ph.EP↗

Extreme Learning Machines for Exoplanet Simulations: A Faster, Lightweight Alternative to Deep Learning

Increasing resolution and coverage of astrophysical and climate data necessitates increasingly sophisticated models, often pushing the limits of computational feasibility. While emulation methods can reduce calculation costs, the neural architectures typically used--optimised via gradient descent--are themselves computationally expensive to train, particularly in terms of data generation requirements. This paper investigates the utility of the Extreme Learning Machine (ELM) as a lightweight, non-gradient-based machine learning algorithm for accelerating complex physical models. We evaluate ELM surrogate models in two test cases with different data structures: (i) sequentially-structured data, and (ii) image-structured data. For test case (i), where the number of samples $N$ >> the dimensionality of input data $d$, ELMs achieve remarkable efficiency, offering a 100,000$\times$ faster training time and a 40$\times$ faster prediction speed compared to a Bi-Directional Recurrent Neural Network (BIRNN), whilst improving upon BIRNN test performance. For test case (ii), characterised by $d >> N$ and image-based inputs, a single ELM was insufficient, but an ensemble of 50 individual ELM predictors achieves comparable accuracy to a benchmark Convolutional Neural Network (CNN), with a 16.4$\times$ reduction in training time, though costing a 6.9$\times$ increase in prediction time. We find different sample efficiency characteristics between the test cases: in test case (i) individual ELMs demonstrate superior sample efficiency, requiring only 0.28% of the training dataset compared to the benchmark BIRNN, while in test case (ii) the ensemble approach requires 78% of the data used by the CNN to achieve comparable results--representing a trade-off between sample efficiency and model complexity.

astro-ph.EP↗

Overview and practical recommendations on using Shapley Values for identifying predictive biomarkers via CATE modeling

In recent years, two parallel research trends have emerged in machine learning, yet their intersections remain largely unexplored. On one hand, there has been a significant increase in literature focused on Individual Treatment Effect (ITE) modeling, particularly targeting the Conditional Average Treatment Effect (CATE) using meta-learner techniques. These approaches often aim to identify causal effects from observational data. On the other hand, the field of Explainable Machine Learning (XML) has gained traction, with various approaches developed to explain complex models and make their predictions more interpretable. A prominent technique in this area is Shapley Additive Explanations (SHAP), which has become mainstream in data science for analyzing supervised learning models. However, there has been limited exploration of SHAP application in identifying predictive biomarkers through CATE models, a crucial aspect in pharmaceutical precision medicine. We address inherent challenges associated with the SHAP concept in multi-stage CATE strategies and introduce a surrogate estimation approach that is agnostic to the choice of CATE strategy, effectively reducing computational burdens in high-dimensional data. Using this approach, we conduct simulation benchmarking to evaluate the ability to accurately identify biomarkers using SHAP values derived from various CATE meta-learners and Causal Forest.

stat.ME↗

Finding Pegasus: Enhancing Unsupervised Anomaly Detection in High-Dimensional Data using a Manifold-Based Approach

Unsupervised machine learning methods are well suited to searching for anomalies at scale but can struggle with the high-dimensional representation of many modern datasets, hence dimensionality reduction (DR) is often performed first. In this paper we analyse unsupervised anomaly detection (AD) from the perspective of the manifold created in DR. We present an idealised illustration, "Finding Pegasus", and a novel formal framework with which we categorise AD methods and their results into "on manifold" and "off manifold". We define these terms and show how they differ. We then use this insight to develop an approach of combining AD methods which significantly boosts AD recall without sacrificing precision in situations employing high DR. When tested on MNIST data, our approach of combining AD methods improves recall by as much as 16 percent compared with simply combining with the best standalone AD method (Isolation Forest), a result which shows great promise for its application to real-world data.

cs.LG↗

Fast Regression of the Tritium Breeding Ratio in Fusion Reactors

The tritium breeding ratio (TBR) is an essential quantity for the design of modern and next-generation D-T fueled nuclear fusion reactors. Representing the ratio between tritium fuel generated in breeding blankets and fuel consumed during reactor runtime, the TBR depends on reactor geometry and material properties in a complex manner. In this work, we explored the training of surrogate models to produce a cheap but high-quality approximation for a Monte Carlo TBR model in use at the UK Atomic Energy Authority. We investigated possibilities for dimensional reduction of its feature space, reviewed 9 families of surrogate models for potential applicability, and performed hyperparameter optimisation. Here we present the performance and scaling properties of these models, the fastest of which, an artificial neural network, demonstrated $R^2=0.985$ and a mean prediction time of $0.898\ μ\mathrm{s}$, representing a relative speedup of $8\cdot 10^6$ with respect to the expensive MC model. We further present a novel adaptive sampling algorithm, Quality-Adaptive Surrogate Sampling, capable of interfacing with any of the individually studied surrogates. Our preliminary testing on a toy TBR theory has demonstrated the efficacy of this algorithm for accelerating the surrogate modelling process.

physics.comp-ph↗

Don't Pay Attention to the Noise: Learning Self-supervised Representations of Light Curves with a Denoising Time Series Transformer

Astrophysical light curves are particularly challenging data objects due to the intensity and variety of noise contaminating them. Yet, despite the astronomical volumes of light curves available, the majority of algorithms used to process them are still operating on a per-sample basis. To remedy this, we propose a simple Transformer model -- called Denoising Time Series Transformer (DTST) -- and show that it excels at removing the noise and outliers in datasets of time series when trained with a masked objective, even when no clean targets are available. Moreover, the use of self-attention enables rich and illustrative queries into the learned representations. We present experiments on real stellar light curves from the Transiting Exoplanet Space Satellite (TESS), showing advantages of our approach compared to traditional denoising techniques.

astro-ph.IM↗

ESA-Ariel Data Challenge NeurIPS 2022: Inferring Physical Properties of Exoplanets From Next-Generation Telescopes

The study of extra-solar planets, or simply, exoplanets, planets outside our own Solar System, is fundamentally a grand quest to understand our place in the Universe. Discoveries in the last two decades have re-defined our understanding of planets, and helped us comprehend the uniqueness of our very own Earth. In recent years the focus has shifted from planet detection to planet characterisation, where key planetary properties are inferred from telescope observations using Monte Carlo-based methods. However, the efficiency of sampling-based methodologies is put under strain by the high-resolution observational data from next generation telescopes, such as the James Webb Space Telescope and the Ariel Space Mission. We are delighted to announce the acceptance of the Ariel ML Data Challenge 2022 as part of the NeurIPS competition track. The goal of this challenge is to identify a reliable and scalable method to perform planetary characterisation. Depending on the chosen track, participants are tasked to provide either quartile estimates or the approximate distribution of key planetary properties. To this end, a synthetic spectroscopic dataset has been generated from the official simulators for the ESA Ariel Space Mission. The aims of the competition are three-fold. 1) To offer a challenging application for comparing and advancing conditional density estimation methods. 2) To provide a valuable contribution towards reliable and efficient analysis of spectroscopic data, enabling astronomers to build a better picture of planetary demographics, and 3) To promote the interaction between ML and exoplanetary science. The competition is open from 15th June and will run until early October, participants of all skill levels are more than welcomed!

astro-ph.EP↗

Peeking inside the Black Box: Interpreting Deep Learning Models for Exoplanet Atmospheric Retrievals

Deep learning algorithms are growing in popularity in the field of exoplanetary science due to their ability to model highly non-linear relations and solve interesting problems in a data-driven manner. Several works have attempted to perform fast retrievals of atmospheric parameters with the use of machine learning algorithms like deep neural networks (DNNs). Yet, despite their high predictive power, DNNs are also infamous for being 'black boxes'. It is their apparent lack of explainability that makes the astrophysics community reluctant to adopt them. What are their predictions based on? How confident should we be in them? When are they wrong and how wrong can they be? In this work, we present a number of general evaluation methodologies that can be applied to any trained model and answer questions like these. In particular, we train three different popular DNN architectures to retrieve atmospheric parameters from exoplanet spectra and show that all three achieve good predictive performance. We then present an extensive analysis of the predictions of DNNs, which can inform us - among other things - of the credibility limits for atmospheric parameters for a given instrument and model. Finally, we perform a perturbation-based sensitivity analysis to identify to which features of the spectrum the outcome of the retrieval is most sensitive. We conclude that for different molecules, the wavelength ranges to which the DNN's predictions are most sensitive, indeed coincide with their characteristic absorption regions. The methodologies presented in this work help to improve the evaluation of DNNs and to grant interpretability to their predictions.

astro-ph.EP↗

PyLightcurve-torch: a transit modelling package for deep learning applications in PyTorch

We present a new open source python package, based on PyLightcurve and PyTorch, tailored for efficient computation and automatic differentiation of exoplanetary transits. The classes and functions implemented are fully vectorised, natively GPU-compatible and differentiable with respect to the stellar and planetary parameters. This makes PyLightcurve-torch suitable for traditional forward computation of transits, but also extends the range of possible applications with inference and optimisation algorithms requiring access to the gradients of the physical model. This endeavour is aimed at fostering the use of deep learning in exoplanets research, motivated by an ever increasing amount of stellar light curves data and various incentives for the improvement of detection and characterisation techniques.

astro-ph.EP↗

Lessons Learned from the 1st ARIEL Machine Learning Challenge: Correcting Transiting Exoplanet Light Curves for Stellar Spots

The last decade has witnessed a rapid growth of the field of exoplanet discovery and characterisation. However, several big challenges remain, many of which could be addressed using machine learning methodology. For instance, the most prolific method for detecting exoplanets and inferring several of their characteristics, transit photometry, is very sensitive to the presence of stellar spots. The current practice in the literature is to identify the effects of spots visually and correct for them manually or discard the affected data. This paper explores a first step towards fully automating the efficient and precise derivation of transit depths from transit light curves in the presence of stellar spots. The methods and results we present were obtained in the context of the 1st Machine Learning Challenge organized for the European Space Agency's upcoming Ariel mission. We first present the problem, the simulated Ariel-like data and outline the Challenge while identifying best practices for organizing similar challenges in the future. Finally, we present the solutions obtained by the top-5 winning teams, provide their code and discuss their implications. Successful solutions either construct highly non-linear (w.r.t. the raw data) models with minimal preprocessing -deep neural networks and ensemble methods- or amount to obtaining meaningful statistics from the light curves, constructing linear models on which yields comparably good predictive performance.

astro-ph.IM↗

Inferring Causal Direction from Observational Data: A Complexity Approach

At the heart of causal structure learning from observational data lies a deceivingly simple question: given two statistically dependent random variables, which one has a causal effect on the other? This is impossible to answer using statistical dependence testing alone and requires that we make additional assumptions. We propose several fast and simple criteria for distinguishing cause and effect in pairs of discrete or continuous random variables. The intuition behind them is that predicting the effect variable using the cause variable should be `simpler' than the reverse -- different notions of `simplicity' giving rise to different criteria. We demonstrate the accuracy of the criteria on synthetic data generated under a broad family of causal mechanisms and types of noise.

cs.LG↗

Margin Maximization as Lossless Maximal Compression

The ultimate goal of a supervised learning algorithm is to produce models constructed on the training data that can generalize well to new examples. In classification, functional margin maximization -- correctly classifying as many training examples as possible with maximal confidence --has been known to construct models with good generalization guarantees. This work gives an information-theoretic interpretation of a margin maximizing model on a noiseless training dataset as one that achieves lossless maximal compression of said dataset -- i.e. extracts from the features all the useful information for predicting the label and no more. The connection offers new insights on generalization in supervised machine learning, showing margin maximization as a special case (that of classification) of a more general principle and explains the success and potential limitations of popular learning algorithms like gradient boosting. We support our observations with theoretical arguments and empirical evidence and identify interesting directions for future work.

cs.LG↗

Pushing the Limits of Exoplanet Discovery via Direct Imaging with Deep Learning

Further advances in exoplanet detection and characterisation require sampling a diverse population of extrasolar planets. One technique to detect these distant worlds is through the direct detection of their thermal emission. The so-called direct imaging technique, is suitable for observing young planets far from their star. These are very low signal-to-noise-ratio (SNR) measurements and limited ground truth hinders the use of supervised learning approaches. In this paper, we combine deep generative and discriminative models to bypass the issues arising when directly training on real data. We use a Generative Adversarial Network to obtain a suitable dataset for training Convolutional Neural Network classifiers to detect and locate planets across a wide range of SNRs. Tested on artificial data, our detectors exhibit good predictive performance and robustness across SNRs. To demonstrate the limits of the detectors, we provide maps of the precision and recall of the model per pixel of the input image. On real data, the models can re-confirm bright source detections.

astro-ph.EP↗

Better Boosting with Bandits for Online Learning

Probability estimates generated by boosting ensembles are poorly calibrated because of the margin maximization nature of the algorithm. The outputs of the ensemble need to be properly calibrated before they can be used as probability estimates. In this work, we demonstrate that online boosting is also prone to producing distorted probability estimates. In batch learning, calibration is achieved by reserving part of the training data for training the calibrator function. In the online setting, a decision needs to be made on each round: shall the new example(s) be used to update the parameters of the ensemble or those of the calibrator. We proceed to resolve this decision with the aid of bandit optimization algorithms. We demonstrate superior performance to uncalibrated and naively-calibrated on-line boosting ensembles in terms of probability estimation. Our proposed mechanism can be easily adapted to other tasks(e.g. cost-sensitive classification) and is robust to the choice of hyperparameters of both the calibrator and the ensemble.

cs.LG↗