TY - RPRT TI - PIRA: Preference-Oriented Instruction-Tuned Reward Models with Dual Aggregation AU - Yongfu Xue PY - 2026 UR - https://arxiv.org/abs/2511.20668 ID - 2511.20668 ER -