TY - RPRT TI - Human Alignment of Large Language Models through Online Preference Optimisation AU - Daniele Calandriello AU - Daniel Guo AU - Remi Munos AU - Mark Rowland AU - Yunhao Tang AU - Bernardo Avila Pires AU - Pierre Harvey Richemond AU - Charline Le Lan AU - Michal Valko AU - Tianqi Liu AU - Rishabh Joshi AU - Zeyu Zheng AU - Bilal Piot PY - 2024 UR - https://arxiv.org/abs/2403.08635 ID - 2403.08635 ER -