TY - RPRT TI - Off-policy Maximum Entropy Reinforcement Learning : Soft Actor-Critic with Advantage Weighted Mixture Policy(SAC-AWMP) AU - Zhimin Hou AU - Kuangen Zhang AU - Yi Wan AU - Dongyu Li AU - Chenglong Fu AU - Haoyong Yu PY - 2020 UR - https://arxiv.org/abs/2002.02829 ID - 2002.02829 ER -