TY - RPRT TI - Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces AU - Chenchen Zhang PY - 2026 UR - https://arxiv.org/abs/2605.02801 ID - 2605.02801 ER -