TY - RPRT TI - Smooth Learning with Hard Constraints via Legendre-Regularized Policies AU - Zikun Lin AU - Rui Chen AU - Yijie Wang PY - 2026 UR - https://arxiv.org/abs/2607.24007 ID - 2607.24007 ER -