TY - RPRT TI - Learning the Value Systems of Agents with Preference-based and Inverse Reinforcement Learning AU - Andrés Holgado-Sánchez AU - Holger Billhardt AU - Alberto Fernández AU - Sascha Ossowski PY - 2026 DO - 10.1007/s10458-026-09732-0 UR - https://arxiv.org/abs/2602.04518 ID - 2602.04518 ER -