TY - RPRT TI - Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game AU - Barna Pásztor AU - Thomas Kleine Buening AU - Andreas Krause PY - 2025 UR - https://arxiv.org/abs/2512.16626 ID - 2512.16626 ER -