TY - RPRT TI - Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor AU - Tuomas Haarnoja AU - Aurick Zhou AU - Pieter Abbeel AU - Sergey Levine PY - 2018 UR - https://arxiv.org/abs/1801.01290 ID - 1801.01290 ER -