TY - RPRT TI - Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models AU - Alessio Galatolo AU - Zhenbang Dai AU - Katie Winkle AU - Meriem Beloucif PY - 2025 UR - https://arxiv.org/abs/2503.03460 ID - 2503.03460 ER -