TY - RPRT TI - Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models AU - Xin Zhou AU - Yiwen Guo AU - Ruotian Ma AU - Tao Gui AU - Qi Zhang AU - Xuanjing Huang PY - 2025 UR - https://arxiv.org/abs/2502.08922 ID - 2502.08922 ER -