TY - RPRT TI - Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models AU - Xuanchen Li AU - Haitao Li AU - Yujia Zhou AU - Qingyi Pan AU - Heng Wang AU - Yiqun Liu AU - Min Zhang AU - Qingyao Ai PY - 2026 UR - https://arxiv.org/abs/2608.09109 ID - 2608.09109 ER -