TY - RPRT TI - The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values AU - Hannah Rose Kirk AU - Andrew M. Bean AU - Bertie Vidgen AU - Paul Röttger AU - Scott A. Hale PY - 2023 UR - https://arxiv.org/abs/2310.07629 ID - 2310.07629 ER -