TY - RPRT TI - ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning AU - Daiki E. Matsunaga AU - Junho Na AU - Tri Wahyu Guntara AU - Scott Sanner AU - Pascal Poupart AU - Jongmin Lee AU - Kee-Eung Kim PY - 2026 UR - https://arxiv.org/abs/2606.30072 ID - 2606.30072 ER -