arXiv · 2508.13661
Communication-Enhanced Tutoring for Efficient Decentralized Multi-Agent Reinforcement Learning
Abstract
Centralized Training with Decentralized Execution (CTDE) is the dominant paradigm in multi-agent reinforcement learning (MARL), enabling agents to act independently at test time while leveraging additional information during training. However, the most prominent methods within CTDE, based on value decomposition, are limited in learning efficiency and final performance by partial observability in both training and execution. To overcome this limitation, in this work, we propose the framework of tutoring: In training, the agents share information in their latent space to develop well-informed policies that achieve strong performance. Then, to recover decentralized execution, these policies concurrently adjust to anticipate lack of communication, and they are distilled into counterparts that rely solely on local observations. We demonstrate the effectiveness of our approach on Hallway, which, to the best of our knowledge, has not been solved before without test-time communication, SMAC under settings more difficult than the standard ones, and SMACv2.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Maciej Wojtala, Bogusz Stefańczyk, Dominik Bogucki, Łukasz Lepak, Paweł Wawrzyński. 2025-08-19. Communication-Enhanced Tutoring for Efficient Decentralized Multi-Agent Reinforcement Learning. https://arxiv.org/abs/2508.13661
Cite the original work for its findings. Save a collection to share your selection of sources.