TY - RPRT TI - Test-time reward-guided alignment of language models by importance sampling on pre-logit space AU - Sekitoshi Kanai AU - Tsukasa Yoshida AU - Hiroshi Takahashi AU - Haru Kuroki AU - Kazumune Hashimoto PY - 2026 UR - https://arxiv.org/abs/2510.26219 ID - 2510.26219 ER -