TY - RPRT TI - Convergence of a Human-in-the-Loop Policy-Gradient Algorithm With Eligibility Trace Under Reward, Policy, and Advantage Feedback AU - Ishaan Shah AU - David Halpern AU - Kavosh Asadi AU - Michael L. Littman PY - 2021 UR - https://arxiv.org/abs/2109.07054 ID - 2109.07054 ER -