TY - RPRT TI - Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection AU - Kyungjae Lee AU - Dasol Hwang AU - Sunghyun Park AU - Youngsoo Jang AU - Moontae Lee PY - 2024 UR - https://arxiv.org/abs/2403.14238 ID - 2403.14238 ER -