TY - RPRT TI - Efficient Offline Reinforcement Learning: First Imitate, then Improve AU - Adam Jelley AU - Trevor McInroe AU - Sam Devlin AU - Amos Storkey PY - 2025 UR - https://arxiv.org/abs/2406.13376 ID - 2406.13376 ER -