SearcharxivSearch

arXiv · 2606.10771

On-sky demonstration of reinforcement learning for adaptive optics control

Abstract

Reinforcement learning (RL)-based algorithms have recently emerged as a promising approach for adaptive optics (AO) control. In simulations and laboratory experiments, they have demonstrated robustness to real-world effects such as photon and detector noise, misregistration, vibrations, and rapid variations in seeing conditions. However, their performance has not yet been validated on sky. We report the first on-sky demonstration of a reinforcement learning controller for adaptive optics, named Policy Optimization for AO (PO4AO). We further analyze its on-sky behavior and identify directions for improving the algorithm and its implementation.PO4AO was implemented and deployed on the Papyrus adaptive optics system installed at the Coud\'e focus of the 1.52 m telescope (T152) at the OHP. A Python-based implementation was interfaced with the existing real-time controller (DAO RTC) via shared-memory buffers. The performance of PO4AO was compared to that of a standard integrator controller over several nights, covering a range of flux levels and atmospheric conditions. PO4AO consistently outperformed the standard integrator in all tested configurations. The controller successfully learned and compensated for vibration patterns and demonstrated strong robustness to measurement noise. Once tuned for Papyrus, PO4AO operated in a turnkey fashion, using a single set of hyperparameters across varying observing conditions and science targets. These performance gains were achieved despite a non-optimized Python implementation introducing approximately $750\,\mu\text{s}$ of additional latency, along with control jitter and occasional frame drops. When properly implemented and optimized, PO4AO constitutes a robust and high-performance turnkey controller for single-conjugate adaptive optics systems, paving the way for broader adoption of reinforcement learning strategies in on-sky AO operations.

Explore related subjects

Keep this discovery

BibTeXRIS

Jalo Nousiainen, Vincent Chambouleyron, Benoit Neichel, Sylvain Cetre, Jean-Francois Sauvage, Angelie Alagao, Markus Kasper, Jonathan Dray, Romain Fetick, Byron Engler. 2026-06-09. On-sky demonstration of reinforcement learning for adaptive optics control. https://arxiv.org/abs/2606.10771

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The EDD Radio Astronomy Backend Framework

Modern digital radio astronomy receivers produce increasingly wide-bandwidth, high bit-rate data streams that necessitate the development of flexible, scalable, and maintainable backend processing and recording systems. Historically, such backend instrumentation has been tightly coupled to telescope observing modes, limiting reuse between observatories and science cases. We present the Effelsberg Direct Digitisation (EDD) backend framework, a software-defined architecture for constructing real-time radio astronomy backends on commodity off-the-shelf computing infrastructure. We describe its design, implementation, supported observing modes, and operational deployments. EDD separates a common core framework from plugin-provided observing capabilities. The core provides orchestration, telescope interfaces, pipeline lifecycle management, monitoring, and deployment tooling, while plugins implement processing pipelines for specific observing modes. The framework is designed to support both single-dish and interferometric instruments through site-specific configuration and plugin selection. EDD currently supports spectroscopy and spectropolarimetry, pulsar timing and searching, baseband recording, very long baseline interferometry, correlation, and beamforming. Operational deployments include the Effelsberg 100-m telescope, the SKA-MPI prototype dish, the Thai National Radio Telescope, and the ARGOS interferometric prototype array. By separating common services, observing-mode plugins, and site-specific configuration, it allows backend capabilities to be deployed across heterogeneous telescope environments and provides a community resource for broadband radio astronomy instrumentation.

astro-ph.IM

Bayesian Superiority in On/Off analysis

We present a detailed comparison of Bayesian criteria with three non-informative priors - flat, Jeffreys, and scale-invariant - for testing a signal against an unknown background and compare them with the classical frequentist Li-Ma approach in the On/Off problem. We perform Monte Carlo simulations for various background levels and evaluate the Li-Ma and Bayesian criteria by their Type I error rates. We then simulate a nonzero signal and compare the criteria in terms of Type II error rates. We find that the Bayesian criterion with the Jeffreys prior yields lower Type I and Type II error rates than the Li-Ma criterion. In addition, we show that the Bayesian criteria are more robust than the Li-Ma criterion when the background distribution is overdispersed relative to the Poisson distribution.

astro-ph.IM

An RFSoC-based Backend and Timing System for the Balloon-borne Very Long Baseline Interferometry Experiment

We present the design and performance characterization of the digital backend and precision-timing system for the Balloon-borne Very Long Baseline Interferometry Experiment (BVEX), a pathfinder for high-frequency stratospheric VLBI at 22 GHz. The backend uses one of the four 14-bit analog-to-digital converter inputs on an AMD-Xilinx RFSoC 4x2. Although the converters support sampling rates up to 5 GSPS, the flight configuration digitizes the 2-4 GHz intermediate frequency at 4.096 GSPS. CASPER firmware provides both a high-resolution spectrometer for pointing and receiver verification, and a VLBI acquisition chain with two-bit requantization that records at a rate of about 8.2 Gbps. The timestamped data packets are sent over 100 Gigabit Ethernet (GbE) to a 16 TB NVMe array in a storage computer that draws approximately 70-80 W. The timing chain uses a Rakon oven-controlled crystal oscillator as a timing reference while a time-interval counter measures its drift relative to a GPS reference with approximately 60 ps resolution. This is the first deployment of an RFSoC-based VLBI backend and precision-timing system on a stratospheric balloon. Ground tests validated the backend, spectrometer, and timing chain. The August 2025 CSA STRATOS flight ended before reaching the target float altitude because of a balloon failure, and as a result no science observations were obtained. For the planned 2027 reflight, we are developing a conduction-cooled data storage computer with 24 TB of NVMe capacity and a direct data path from the 100 GbE interface to the NVMe array.

astro-ph.IM