TY - RPRT TI - Revisiting the Learning Objectives of Vision-Language Reward Models AU - Simon Roy AU - Samuel Barbeau AU - Giovanni Beltrame AU - Christian Desrosiers AU - Nicolas Thome PY - 2025 UR - https://arxiv.org/abs/2512.20675 ID - 2512.20675 ER -