TY - RPRT TI - The History and Risks of Reinforcement Learning and Human Feedback AU - Nathan Lambert AU - Thomas Krendl Gilbert AU - Tom Zick PY - 2023 UR - https://arxiv.org/abs/2310.13595 ID - 2310.13595 ER -