TY - RPRT TI - Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology AU - Jian Xiong AU - Jingbo Zhou AU - Zihan Zhou AU - Yixiong Xiao AU - Le Zhang AU - Jingyong Ye AU - Rui Qian AU - Yang Zhou AU - Dejing Dou PY - 2026 UR - https://arxiv.org/abs/2601.22474 ID - 2601.22474 ER -