TY - RPRT TI - Learning the Value Systems of Societies with Preference-based Multi-objective Reinforcement Learning AU - Andrés Holgado-Sánchez AU - Peter Vamplew AU - Richard Dazeley AU - Sascha Ossowski AU - Holger Billhardt PY - 2026 UR - https://arxiv.org/abs/2602.08835 ID - 2602.08835 ER -