TY - RPRT TI - Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data AU - Jeonghye Kim AU - Yongjae Shin AU - Whiyoung Jung AU - Sunghoon Hong AU - Deunsol Yoon AU - Youngchul Sung AU - Kanghoon Lee AU - Woohyung Lim PY - 2025 UR - https://arxiv.org/abs/2507.08761 ID - 2507.08761 ER -