TY - RPRT TI - Strategyproof Reinforcement Learning from Human Feedback AU - Thomas Kleine Buening AU - Jiarui Gan AU - Debmalya Mandal AU - Marta Kwiatkowska PY - 2025 UR - https://arxiv.org/abs/2503.09561 ID - 2503.09561 ER -