arXiv · 2507.18992
Reinforcement Learning via Conservative Agent for Environments with Random Delays
Abstract
Real-world reinforcement learning applications are often hindered by delayed feedback from environments, which violates the Markov assumption and introduces significant challenges. Although numerous delay-compensating methods have been proposed for environments with constant delays, environments with random delays remain largely unexplored due to their inherent variability and unpredictability. In this study, we propose a simple yet robust agent for decision-making under random delays, termed the conservative agent, which reformulates the random-delay environment into its constant-delay equivalent. This transformation enables any state-of-the-art constant-delay method to be directly extended to the random-delay environments without modifying the algorithmic structure or sacrificing performance. We evaluate the conservative agent-based algorithm on continuous control tasks, and empirical results demonstrate that it significantly outperforms existing baseline algorithms in terms of asymptotic performance and sample efficiency.
Explore related subjects
Keep this discovery
Jongsoo Lee, Jangwon Kim, Jiseok Jeong, Soohee Han. 2025-07-25. Reinforcement Learning via Conservative Agent for Environments with Random Delays. https://arxiv.org/abs/2507.18992
Cite the original work for its findings. Save a collection to share your selection of sources.