SearcharxivSearch

arXiv subjects

Ulrich Armel Mbou Sob

Publications and source records attributed to Ulrich Armel Mbou Sob.

5 recordsLinked to original sources

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

Generalising to unseen tasks remains a fundamental challenge in offline multi-agent reinforcement learning (MARL). In this work, we present a principled analysis of zero-shot task generalisation in the offline setting and conduct an extensive empirical investigation into the scaling behaviour governing task diversity, dataset size, and network capacity. To facilitate this study, we extend offline sequence modelling architectures to handle multi-task observation and action spaces alongside variable agent counts across tasks. Our primary finding is that scaling task diversity---rather than sheer dataset size is the dominant factor in achieving robust zero-shot transfer. Through large-scale experiments across four challenging environments (Connector, RWARE, SMAX, and LBF), we demonstrate that our multi-task approach achieves a mean improvement of 3.2x on held-out test tasks compared to single-task models and consistently outperforms strong behaviour cloning baselines. These results suggest that the development of generalisable MARL agents should prioritise the diversity of the training distribution with varying numbers of agents, providing a roadmap for scaling offline MARL effectively.

cs.LG

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations, removing the need for hand-crafted reward functions. Despite impressive success in achieving viable self-supervised learning in RL, all existing CRL algorithms rely on off-policy optimisation and are mostly constrained to continuous action spaces, with little research invested in discrete environments. This leaves CRL disconnected from widely used and effective, modern on-policy training pipelines adopted across both single-agent and multi-agent RL in continuous and discrete environments. To establish a first connection, we introduce Contrastive Proximal Policy Optimisation (CPPO). CPPO is an on-policy contrastive RL algorithm that derives policy advantages directly from contrastive Q-values and optimises them via the standard PPO objective, without requiring a reward function or a replay buffer. We evaluate CPPO across continuous and discrete, single-agent and cooperative multi-agent tasks. Whilst the existence of an on-policy approach is inherently useful, we observe that \textbf{CPPO not only significantly outperforms the previous CRL baselines in 14 out of 18 tasks, but also matches or exceeds PPO's performance, which uses hand-crafted dense rewards, in 12 out of the 18 tasks tested.}

cs.LG

The Hydrogen Intensity and Real-time Analysis eXperiment: 256-Element Array Status and Overview

The Hydrogen Intensity and Real-time Analysis eXperiment (HIRAX) is a radio interferometer array currently in development, with an initial 256-element array to be deployed at the South African Radio Astronomy Observatory (SARAO) Square Kilometer Array (SKA) site in South Africa. Each of the 6m, $f/0.23$ dishes will be instrumented with dual-polarisation feeds operating over a frequency range of 400-800 MHz. Through intensity mapping of the 21 cm emission line of neutral hydrogen, HIRAX will provide a cosmological survey of the distribution of large-scale structure over the redshift range of $0.775 < z < 2.55$ over $\sim$15,000 square degrees of the southern sky. The statistical power of such a survey is sufficient to produce $\sim$7 percent constraints on the dark energy equation of state parameter when combined with measurements from the Planck satellite. Additionally, HIRAX will provide a highly competitive platform for radio transient and HI absorber science while enabling a multitude of cross-correlation studies. In this paper, we describe the science goals of the experiment, overview of the design and status of the sub-components of the telescope system, and describe the expected performance of the initial 256-element array as well as the planned future expansion to the final, 1024-element array.

astro-ph.IM

Solution intervals considered harmful: on the optimality of radio interferometric gain solutions

Solution intervals are often used to improve the signal-to-noise ratio during radio interferometric gain calibration. This work investigates how factors such as the noise level, intrinsic gain variability, degree of model incompleteness, and the presence of radio frequency interference impact the selection of solution intervals for calibration. We perform different interferometric simulations to demonstrate how these factors, in combination with the choice of solution intervals, affect calibration and imaging outputs and discuss practical guidelines for choosing optimal solution intervals. Furthermore, we present an algorithm capable of automatically selecting suitable solution intervals during calibration. By applying the algorithm to both simulated and real data, we show that it can successfully choose solution intervals that strike a good balance between capturing intrinsic gain variability and not fitting noise as long as the data are not too inhomogeneously flagged. Furthermore, we elaborate on several practical aspects that emphasize the need to develop regularised calibration algorithms that do not require solution intervals.

astro-ph.IM

Radio Interferometric Calibration Using a Complex Student's t-distribution and Wirtinger Derivatives

Radio interferometric gain calibration can be biased by incomplete sky models and radio frequency interference, resulting in calibration artefacts that can restrict the dynamic range of the resulting images. It has been suggested that calibration algorithms employing heavy-tailed likelihood functions are less susceptible to this due to their robustness against outliers in the data. We present an algorithm based on a Student's t-distribution which leverages the framework of complex optimisation and Wirtinger calculus for efficient and robust interferometric gain calibration. We integrate this algorithm as an option in the newly released calibration software package, CubiCal. We demonstrate that the algorithm can mitigate some of the biases introduced by incomplete sky models and radio frequency interference by applying it to both simulated and real data. Our results show significant improvements compared to a conventional least-squares solver which assumes a Gaussian likelihood function. Furthermore, we provide some insight into why the algorithm outperforms the conventional solver, and discuss specific scenarios (for both direction-independent and direction-dependent self-calibration) where this is expected to be the case.

astro-ph.IM