TY - RPRT TI - A Minimaximalist Approach to Reinforcement Learning from Human Feedback AU - Gokul Swamy AU - Christoph Dann AU - Rahul Kidambi AU - Zhiwei Steven Wu AU - Alekh Agarwal PY - 2024 UR - https://arxiv.org/abs/2401.04056 ID - 2401.04056 ER -