TY - RPRT TI - SOAP-RL: Sequential Option Advantage Propagation for Reinforcement Learning in POMDP Environments AU - Shu Ishida AU - João F. Henriques PY - 2024 UR - https://arxiv.org/abs/2407.18913 ID - 2407.18913 ER -