TY - RPRT TI - On the Convergence and Stability of Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning, and Online Decision Transformers AU - Miroslav Štrupl AU - Oleg Szehr AU - Francesco Faccio AU - Dylan R. Ashley AU - Rupesh Kumar Srivastava AU - Jürgen Schmidhuber PY - 2025 UR - https://arxiv.org/abs/2502.05672 ID - 2502.05672 ER -