Searcharxiv⌕ Search

arXiv subjects

Jalo Nousiainen

Publications and source records attributed to Jalo Nousiainen.

8 recordsLinked to original sources

Exploring reinforcement learning to enhance focal-plane wavefront control for vortex coronagraphs

High Contrast Imaging (HCI) on ground-based telescopes suffers from phase aberrations on the observed wavefront caused by atmospheric turbulence. Adaptive Optics (AO) systems are adept at correcting these aberrations, but fall short in the correction of non-common path aberrations (NCPAs). NCPAs arise because the wavefront sensor (WFS) measures and corrects a wavefront that is different from that affecting the science images, thus requiring additional intervention. This work makes use of focal-plane wavefront sensing and reinforcement learning (RL) to address the wavefront aberrations caused by NCPAs. The PO4NCPA algorithm utilizes sequential phase diversity to address phase ambiguities and is tested on a simulation designed for the Mid-infrared ELT Imager and Spectrograph (METIS) instrument. In this paper, we present the performance of PO4NCPA with scalar and vector vortex coronagraphs to demonstrate its flexibility.

astro-ph.IM↗

On-sky demonstration of reinforcement learning for adaptive optics control

Reinforcement learning (RL)-based algorithms have recently emerged as a promising approach for adaptive optics (AO) control. In simulations and laboratory experiments, they have demonstrated robustness to real-world effects such as photon and detector noise, misregistration, vibrations, and rapid variations in seeing conditions. However, their performance has not yet been validated on sky. We report the first on-sky demonstration of a reinforcement learning controller for adaptive optics, named Policy Optimization for AO (PO4AO). We further analyze its on-sky behavior and identify directions for improving the algorithm and its implementation.PO4AO was implemented and deployed on the Papyrus adaptive optics system installed at the Coudé focus of the 1.52 m telescope (T152) at the OHP. A Python-based implementation was interfaced with the existing real-time controller (DAO RTC) via shared-memory buffers. The performance of PO4AO was compared to that of a standard integrator controller over several nights, covering a range of flux levels and atmospheric conditions. PO4AO consistently outperformed the standard integrator in all tested configurations. The controller successfully learned and compensated for vibration patterns and demonstrated strong robustness to measurement noise. Once tuned for Papyrus, PO4AO operated in a turnkey fashion, using a single set of hyperparameters across varying observing conditions and science targets. These performance gains were achieved despite a non-optimized Python implementation introducing approximately $750\,μ\text{s}$ of additional latency, along with control jitter and occasional frame drops. When properly implemented and optimized, PO4AO constitutes a robust and high-performance turnkey controller for single-conjugate adaptive optics systems, paving the way for broader adoption of reinforcement learning strategies in on-sky AO operations.

astro-ph.IM↗

Focal plane wavefront control with model-based reinforcement learning

The direct imaging of potentially habitable exoplanets is one prime science case for high-contrast imaging instruments on extremely large telescopes. Most such exoplanets orbit close to their host stars, where their observation is limited by fast-moving atmospheric speckles and quasi-static non-common-path aberrations (NCPA). Conventional NCPA correction methods often use mechanical mirror probes, which compromise performance during operation. This work presents machine-learning-based NCPA control methods that automatically detect and correct both dynamic and static NCPA errors by leveraging sequential phase diversity. We extend previous work in reinforcement learning for AO to focal plane control. A new model-based RL algorithm, Policy Optimization for NCPAs (PO4NCPA), interprets the focal-plane image as input data and, through sequential phase diversity, determines phase corrections that optimize both non-coronagraphic and post-coronagraphic PSFs without prior system knowledge. Further, we demonstrate the effectiveness of this approach by numerically simulating static NCPA errors on a ground-based telescope and an infrared imager affected by water-vapor-induced seeing (dynamic NCPAs). Simulations show that PO4NCPA robustly compensates static and dynamic NCPAs. In static cases, it achieves near-optimal focal-plane light suppression with a coronagraph and near-optimal Strehl without one. With dynamics NCPA, it matches the performance of the modal least-squares reconstruction combined with a 1-step delay integrator in these metrics. The method remains effective for the ELT pupil, vector vortex coronagraph, and under photon and background noise. PO4NCPA is model-free and can be directly applied to standard imaging as well as to any coronagraph. Its sub-millisecond inference times and performance also make it suitable for real-time low-order correction of atmospheric turbulence beyond HCI.

astro-ph.IM↗

The GPU-based High-order adaptive OpticS Testbench

The GPU-based High-order adaptive OpticS Testbench (GHOST) at the European Southern Observatory (ESO) is a new 2-stage extreme adaptive optics (XAO) testbench at ESO. The GHOST is designed to investigate and evaluate new control methods (machine learning, predictive control) for XAO which will be required for instruments such as the Planetary Camera and Spectrograph of ESOs Extremely Large Telescope. The first stage corrections are performed in simulation, with the residual wavefront error at each iteration saved. The residual wavefront errors from the first stage are then injected into the GHOST using a spatial light modulator. The second stage correction is made with a Boston Michromachines Corporation 492 actuator deformable mirror and a pyramid wavefront sensor. The flexibility of the bench also opens it up to other applications, one such application is investigating the flip-flop modulation method for the pyramid wavefront sensor.

astro-ph.IM↗

The power of prediction: spatiotemporal Gaussian process modeling for predictive control in slope-based wavefront sensing

Time-delay error is a significant error source in adaptive optics (AO) systems. It arises from the latency between sensing the wavefront and applying the correction. Predictive control algorithms reduce the time-delay error, providing significant performance gains, especially for high-contrast imaging. However, the predictive controller's performance depends on factors such as the WFS type, the measurement noise, the AO system's geometry, and the atmospheric conditions. This work studies the limits of prediction under different imaging conditions through spatiotemporal Gaussian process models. The method provides a predictive reconstructor that is optimal in the least-squares sense, conditioned on the fixed times series of WFS data and our knowledge of the atmosphere. We demonstrate that knowledge is power in predictive AO control. With an SHS-based extreme AO instrument, perfect knowledge of Frozen Flow evolution (wind and Cn2 profile) leads to a reduction of the residual wavefront phase variance up to a factor of 3.5 compared to a non-predictive approach. If there is uncertainty in the profile or evolution models, the gain is more modest. Still, assuming that only effective wind speed is available (without direction) led to reductions in variance by a factor of 2.3. We also study the value of data for predictive filters by computing the experimental utility for different scenarios to answer questions such as: How many past data frames should the prediction filter consider, and is it always most advantageous to use the most recent data? We show that within the scenarios considered, more data consistently increases prediction accuracy. Further, we demonstrate that given a computational limitation on how many past frames we can use, an optimized selection of $n$ past frames leads to a 10-15% additional improvement in RMS over using the n latest consecutive frames of data.

astro-ph.IM↗

Laboratory Experiments of Model-based Reinforcement Learning for Adaptive Optics Control

Direct imaging of Earth-like exoplanets is one of the most prominent scientific drivers of the next generation of ground-based telescopes. Typically, Earth-like exoplanets are located at small angular separations from their host stars, making their detection difficult. Consequently, the adaptive optics (AO) system's control algorithm must be carefully designed to distinguish the exoplanet from the residual light produced by the host star. A new promising avenue of research to improve AO control builds on data-driven control methods such as Reinforcement Learning (RL). RL is an active branch of the machine learning research field, where control of a system is learned through interaction with the environment. Thus, RL can be seen as an automated approach to AO control, where its usage is entirely a turnkey operation. In particular, model-based reinforcement learning (MBRL) has been shown to cope with both temporal and misregistration errors. Similarly, it has been demonstrated to adapt to non-linear wavefront sensing while being efficient in training and execution. In this work, we implement and adapt an RL method called Policy Optimization for AO (PO4AO) to the GHOST test bench at ESO headquarters, where we demonstrate a strong performance of the method in a laboratory environment. Our implementation allows the training to be performed parallel to inference, which is crucial for on-sky operation. In particular, we study the predictive and self-calibrating aspects of the method. The new implementation on GHOST running PyTorch introduces only around 700 microseconds in addition to hardware, pipeline, and Python interface latency. We open-source well-documented code for the implementation and specify the requirements for the RTC pipeline. We also discuss the important hyperparameters of the method, the source of the latency, and the possible paths for a lower latency implementation.

astro-ph.IM↗

Adaptive Optics control using Model-Based Reinforcement Learning

Reinforcement Learning (RL) presents a new approach for controlling Adaptive Optics (AO) systems for Astronomy. It promises to effectively cope with some aspects often hampering AO performance such as temporal delay or calibration errors. We formulate the AO control loop as a model-based RL problem (MBRL) and apply it in numerical simulations to a simple Shack-Hartmann Sensor (SHS) based AO system with 24 resolution elements across the aperture. The simulations show that MBRL controlled AO predicts the temporal evolution of turbulence and adjusts to mis-registration between deformable mirror and SHS which is a typical calibration issue in AO. The method learns continuously on timescales of some seconds and is therefore capable of automatically adjusting to changing conditions.

astro-ph.IM↗

PCS -- A Roadmap for Exoearth Imaging with the ELT

The Planetary Camera and Spectrograph (PCS) for the Extremely Large Telescope (ELT) will be dedicated to detecting and characterising nearby exoplanets with sizes from sub-Neptune to Earth-size in the neighbourhood of the Sun. This goal is achieved by a combination of eXtreme Adaptive Optics (XAO), coronagraphy and spectroscopy. PCS will allow us not only to take images, but also to look for biosignatures such as molecular oxygen in the exoplanets' atmospheres. This article describes the PCS primary science goals, the instrument concept and the research and development activities that will be carried out over the coming years.

astro-ph.IM↗