TY - RPRT TI - HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime AU - Mohamed Sana AU - Nicola Piovesan AU - Antonio De Domenico AU - Fadhel Ayed AU - Haozhe Zhang PY - 2026 UR - https://arxiv.org/abs/2605.30201 ID - 2605.30201 ER -