TY - RPRT TI - Principled Reinforcement Learning with Human Feedback from Pairwise or $K$-wise Comparisons AU - Banghua Zhu AU - Jiantao Jiao AU - Michael I. Jordan PY - 2024 UR - https://arxiv.org/abs/2301.11270 ID - 2301.11270 ER -