TY - RPRT TI - Policy Learning from Large Vision-Language Model Feedback without Reward Modeling AU - Tung M. Luu AU - Donghoon Lee AU - Younghwan Lee AU - Chang D. Yoo PY - 2025 UR - https://arxiv.org/abs/2507.23391 ID - 2507.23391 ER -