TY - RPRT TI - Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs AU - Yifan Zhou AU - Sachin Grover AU - Mohamed El Mistiri AU - Kamalesh Kalirathnam AU - Pratyush Kerhalkar AU - Swaroop Mishra AU - Neelesh Kumar AU - Sanket Gaurav AU - Oya Aran AU - Heni Ben Amor PY - 2025 UR - https://arxiv.org/abs/2511.21928 ID - 2511.21928 ER -